Bold AI Predictions From Cohere Co-founder
Disclaimer: This show is part of our Cohere partnership series. Ivan Zhang, co-founder of Cohere, sits down with Tim Scarfe for a wide-ranging conversation about building enterprise AI from the ground up. Ivan recounts dropping out of the University of Toronto to chase the startup bug, eventually co-founding Cohere with Aidan Gomez (a transformer paper co-author) and Nick Frosst after Google failed to commercialize its own invention -- a textbook innovator's dilemma. The conversation covers Cohere's RAG-first approach to enterprise AI, where models are trained to operate in air-gapped environments relying on external knowledge bases rather than memorized data. A healthcare case study stands out: a simple RAG feature cut doctor preparation time from 30 to 5 minutes. Ivan discusses the practical challenges of GPU allocation in Kubernetes, the tension between consultancy and product, and why the first version of any product will never be right without customer feedback. A surprising detour into competitive gaming reveals Ivan's League of Legends and Elden Ring habits, and the organizational lessons they carry -- communication discipline, not overloading the comms, and situational execution over grand plans. The conversation lands on a striking analogy: debugging LLM context is like understanding human decisions given the information available. "We are all just reasoning engines processing information." The final third covers transformer architecture persistence, system-level optimization over model-only scaling, inference-time computation as a paradigm shift, and the challenge of capturing human thought processes (not just final answers) in training data. Ivan closes with advice for young developers: pair up with AI tools and build things fast. This is the best time to be a developer. --- REFERENCES: company: [00:00:01] Cohere https://cohere.com/ person: [00:01:20] Ivan Zhang https://ivanzhang.ca/ paper: [00:02:40] Attention Is All You Need https://arxiv.org/abs/1706.03762 [00:18:00] Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks https://arxiv.org/abs/2005.11401 [00:35:39] Lets Verify Step by Step https://arxiv.org/abs/2305.20050 [00:39:20] Adaptive Inference-Time Compute https://arxiv.org/abs/2410.02725 [00:43:10] Getting 50% SOTA on ARC-AGI with GPT-4o https://redwoodresearch.substack.com/p/getting-50-sota-on-arc-agi-with-gpt book: [00:03:20] The Innovators Dilemma https://www.amazon.com/Innovators-Dilemma-Technologies-Management-Innovation/dp/1633691780 concept: [00:09:15] Actor Model https://en.wikipedia.org/wiki/Actor_model [00:14:35] Chinese Room Argument https://plato.stanford.edu/entries/chinese-room/ documentation: [00:18:40] Cohere RAG Documentation https://docs.cohere.com/v2/docs/retrieval-augmented-generation-rag --- LINKS: Full Transcript: https://app.rescript.info/share/c4aaa75b8a090bbc48593e47ff39c0f9 Download PDF transcript: https://app.rescript.info/api/public/sessions/63d9341c4891b0c9/pdf https://cohere.com/ https://ivanzhang.ca/ https://x.com/1vnzh REFS: 00:02:40 The Transformer architecture, https://arxiv.org/abs/1706.03762 00:03:22 The Innovator's Dilemma, https://www.amazon.com/Innovators-Dilemma-Technologies-Management-Innovation/dp/1633691780 00:09:15 The actor model, https://en.wikipedia.org/wiki/Actor_model 00:14:35 John Searle's Chinese Room Argument, https://plato.stanford.edu/entries/chinese-room/ 00:18:00 Retrieval-Augmented Generation, https://arxiv.org/abs/2005.11401 00:18:40 Retrieval-Augmented Generation, https://docs.cohere.com/v2/docs/retrieval-augmented-generation-rag 00:35:39 Let’s Verify Step by Step, https://arxiv.org/pdf/2305.20050 00:39:20 Adaptive Inference-Time Compute, https://arxiv.org/abs/2410.02725 00:43:20 Ryan Greenblatt ARC entry, https://redwoodresearch.substack.com/p/getting-50-sota-on-arc-agi-with-gpt