The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI)
Key insights
Books referenced
- Story of Your Life - Ted Chiang - Edwin's all-time favorite short story, about a linguist deciphering an alien language; he rereads it every couple of years and it is the basis for the movie Arrival.
- The Myth of Sisyphus - Albert Camus - Edwin says he can't fully explain why, but the final chapter has always struck him as inspiring.
- Le Ton beau de Marot - Douglas Hofstadter - Edwin's favorite Hofstadter book, ahead of the more famous Godel, Escher, Bach. It translates a single French poem 89 different ways and argues translation isn't robotic, an idea Edwin says maps directly onto how he thinks about data quality.
Media referenced
- Arrival - movie - The film adaptation of Story of Your Life; Edwin loves it as much as the source story.
- Travelers - show - One of Edwin's new favorite TV shows, about people from the future sent back in time to prevent a catastrophe.
- Contact - movie - A movie Edwin recently rewatched; fits his lifelong fascination with scientists deciphering alien communication.
- Dwarkesh Podcast episode with Richard Sutton - podcast - Lenny brings up Sutton's 'bitter lesson' argument that LLMs may be a dead end; Edwin agrees something new beyond LLMs will likely be needed.
Companies
- Surge AI - Edwin's company; trains and evaluates AI models for frontier labs, hit over $1 billion in revenue in under four years with under 100 people, fully bootstrapped and profitable from day one.
- Anthropic - Edwin repeatedly praises Anthropic as the most principled lab on model behavior; also cites Claude's artifacts feature as an underhyped product direction.
- OpenAI - GPT-3's 2020 launch is what convinced Edwin high-quality data was the missing piece and prompted him to start Surge a month later; ChatGPT is cited as an example of sycophantic model behavior.
- Google - One of Edwin's prior employers as a researcher; he uses Google Search's best-of-the-best-versus-worst-of-the-worst ranking problem as an analogy for how Surge measures data quality.
- Meta (Facebook) - A prior employer; Edwin cites Facebook's engagement-optimization era (clickbait, dark patterns) as a cautionary parallel to AI labs now optimizing for engagement.
- Twitter - A prior employer, where Edwin built the well-known soda-versus-pop dialect map.
- xAI - Cited as an example of a model (Grok) with a visibly distinct personality, supporting Edwin's point that models will keep diverging rather than converging.
- DeepMind - Edwin says he was always a huge fan of DeepMind as a research-driven company that stayed scientific even after being acquired; a model he consciously tried to build Surge like.
- Waymo - Edwin rode a Waymo for the first time on a recent SF trip and calls it magical, saying it exceeds the hype.
- Coinbase - Lenny compares Edwin's cross-disciplinary background (math, linguistics, engineering) to Brian Armstrong's origin story as the ideal founder-market fit for Coinbase.
- Vanta - Podcast sponsor; compliance automation platform (SOC 2, ISO 27001).
- WorkOS - Podcast sponsor; provides drop-in enterprise-readiness APIs (SSO, SCIM, RBAC, audit logs) for B2B SaaS.
- Coda - Podcast sponsor; all-in-one workspace Lenny says he uses daily to run his podcast and community.
Techniques and frameworks
- RL environments - Full simulated worlds (a startup's Gmail, Slack, Jira, GitHub, and code base, for example) where a model is given a multi-step goal and rewarded on outcome; designed to expose failures that single-step benchmarks hide.
- Post-training progression: SFT to RLHF to rubrics/verifiers to RL environments - Edwin frames the field's post-training history as a sequence of complementary learning methods, not replacements: supervised fine-tuning (mimicking a master), RLHF (comparing many outputs), rubrics and verifiers (graded feedback), and now RL environments (learning by acting in a simulated world).
- Trajectory-aware evaluation - Grading a model on the full path it took to an answer, not just whether the final answer is correct, since a model can flail through 50 failed attempts, reward-hack, or land on the right number inefficiently.
- Best-of-the-best signal aggregation - Surge tracks thousands of per-worker, per-task signals (keystrokes, response time, review outcomes, downstream model performance) to identify top contributors, framed as analogous to how a search engine both removes spam and surfaces the best page, not just one or the other.
Summary
Edwin Chen is the founder and CEO of Surge AI, the human-data company that trains and evaluates models for essentially every frontier AI lab. The headline framing of the episode is Surge's growth story: over $1 billion in revenue in under four years, fewer than 100 employees, fully bootstrapped, and profitable from day one, achieved without ever posting on LinkedIn, tweeting about the company, or raising VC money. Edwin treats this not as an isolated quirk but as a deliberate rejection of the standard startup playbook: no constant pivoting, no blitzscaling, no chasing valuations, just building the one product only Surge's team could build and letting quality drive word of mouth among researchers who actually cared about data.
Much of the conversation digs into what "data quality" actually means, since Edwin argues almost nobody in the industry defines it seriously. His running example is asking a model to write an eight-line poem about the moon: checking length and topic is trivial, but genuinely high-quality output, poetry with unique imagery and emotional weight, is a much harder and more subjective target. Surge's answer is to track thousands of granular signals per worker and per task (keystrokes, response speed, review outcomes, how outputs actually move model performance) to both filter out the worst contributors and, more importantly, surface the best, an approach Edwin compares to how a search engine both removes spam and ranks the truly excellent page. He credits this depth of data taste, more than raw volume, for why Anthropic's models pulled ahead on coding and writing.
Edwin is sharply critical of how the industry measures progress. He says he doesn't trust benchmarks at all, partly because many contain outright errors, and partly because their clean, well-defined structure is easy for labs to climb without it reflecting real-world usefulness, the reason models can win IMO gold medals while still failing to parse a messy PDF. He's more pointed about crowdsourced leaderboards like LM Arena, where casual voters skim responses for a couple of seconds and reward heavy formatting, emojis, and length over correctness, effectively training labs to produce what he calls "AI slop" optimized for the same instincts that sell tabloids at a grocery store checkout. He connects this to a broader worry about labs optimizing for engagement the way social media once did, warning that sycophantic model behavior (endless praise, endless "helpful" iteration) is the same dynamic that filled social feeds with clickbait.
The second half of the conversation turns to where post-training is heading. Edwin walks through the field's evolution from supervised fine-tuning, to RLHF, to rubrics and verifiers, to what he sees as the current frontier: RL environments, full simulated worlds (a startup's Gmail, Slack, Jira tickets, and code base, for instance) where a model is given a multi-step goal and rewarded on outcome. These environments expose failures that isolated, single-step benchmarks hide, since a model that looks capable at one tool call can fail catastrophically once its actions in step one affect what happens fifty steps later. He also stresses that trajectories matter, not just final answers, since a model can reach the right number through fifty inefficient or accidental attempts, and grading only the endpoint misses what actually needs to be taught.
Underneath the technical discussion runs a consistent philosophical thread: Edwin believes each lab's values will increasingly shape distinct model personalities, using the example of a model that either endlessly polishes an email or tells the user it's good enough and to move on. He frames Surge's deepest job as helping customers define their "dream objective function," comparing it to the difference between raising a child to hit test scores versus raising them to be a genuinely good, happy person: both are legitimate goals, but only one is easy to measure, and most of the industry defaults to the easy one. The episode closes with a personal note on how Edwin's own path (a math and linguistics background at MIT, research roles at Google, Facebook, and Twitter, and a lifelong fascination with deciphering unfamiliar language) shaped Surge's mission, along with a lightning round covering his favorite books, shows, and the soda-versus-pop dialect map that made him briefly internet-famous before he started the company.
Notable Quotes
"We're basically teaching our models to chase dopamine instead of truth." - Edwin Chen
"The easiest way to climb LM Arena is adding crazy bolding, doubling the number of emojis, tripling the length of your model response, even if your model starts hallucinating and getting the answer completely wrong." - Edwin Chen
"Just build the one thing only you could build, the thing that wouldn't exist without the insight and expertise that only you have." - Edwin Chen
"You are your objective function." - Edwin Chen