"Our strong view is that 100 per cent of the data that labs will get value out of will have some human input in the foreseeable future," Alex Ratner said. "But 100 per cent of that data will have to use synthetic and automated approaches to keep up with this complexity."
Demand, he told the reporters, had moved past simpler labeling work. Frontier labs now wanted harder, higher-stakes data to train and evaluate systems that kept growing more capable. The shift was already reshaping what counted as useful training material. That combination—human judgment threaded through every valuable set, automation required to keep pace—was the bet behind the new capital. Snorkel AI had raised $350 million in a Series E. The round valued the company at $3.5 billion.
Four years of research at the Stanford AI Lab, led by co-founder and CEO Alex Ratner and his team, produced the technology that became Snorkel. The company spun out and took commercial form in 2019. By the Series E announcement in 2026 it was seven years old. Its first life was software that automated data labeling. Supervised learning, one of the main ways AI systems are trained, works by showing a model large numbers of examples that already carry the correct answers. Attaching those answers by hand is slow and expensive work; the original product attacked that cost. Only later did Snorkel move from selling labeling tools to delivering finished datasets and the agentic environments in which models train and improve.
In September 2025 Snorkel launched data-as-a-service. It stopped selling tools that helped clients label their own material and began delivering finished products: complete datasets and reinforcement-learning environments. Reinforcement learning trains a model on unanswered tasks, then scores the results so the system can improve; Snorkel builds both the tasks and the simulated settings in which models attempt them. The stack is hybrid. Company software and models generate synthetic data while subject-matter experts design the scenarios, tasks, and grading rubrics. A network of tens of thousands of specialists carries the expert side, with the heaviest demand in coding and steady work in law and medicine. Customers span frontier AI labs, hyperscalers, enterprises, and the U.S. federal government. Because the company sells the datasets and environments rather than human labor, payments to those specialists sit in cost of goods sold. The headline revenue figure is not a pass-through of wages.
Andy Harrison described the arrangement directly: "Snorkel's unique expert-agentic environments create a synergistic relationship between human skills and AI, offering quality data at unprecedented speeds. This capability enables a new era of efficiency and precision in model development." The factory that now produces that data had taken years to reach this form.
"The frontier of AI needs research partners who are at the cutting edge of data science," Ratner said. "Snorkel was designed to be that partner—drawing on over ten years of research and development to be the leading lab for agentic data."
The previous close had been a $100 million Series D in May 2025, struck at a $1.3 billion valuation. Seventeen months separated that mark from the Series E. Capital raised across every round now exceeded $500 million.
Lonne Jaffe, managing director at Insight Partners, said the pace tracked the work itself. "Snorkel's research-grade approach to AI data, environments, and measurement is becoming an increasingly important ingredient in building capable and reliable AI systems – and their demonstrated growth at scale reflects that." Other data firms were posting large top-line numbers of their own. Mercor’s gross annualized revenue has climbed to $2 billion. Handshake reached $1 billion earlier this year. Micro1 scaled to $500 million. Other data companies that position themselves as AI data labs have posted the same kind of climb. The shared cost structure behind those climbs is what investors have to read through. Typical data companies pay 60 to 70 percent of top-line revenue directly to the domain specialists performing the work. Net annual revenue sits substantially lower than the gross headlines that define the set.
Insight Partners and S32 co-led Snorkel’s Series E. "Snorkel's expert-agentic environments create a unique flywheel between human expertise and AI, delivering data at the speed and quality that enables a new era of efficiency and accuracy for model development," said Andy Harrison, CEO and general partner at S32. "Data is becoming more rare, more specialized, more difficult to find," he said. "If you want to train the most frontier, complex and capable models, now you need superior data."
Addition, Lightspeed, Greylock, GV, and Wells Fargo participated, joined by March Capital, Blumberg Capital, Allegis Capital, Frontline, Standard, Third Point Ventures, Factory, Prosperity7, and Walden Catalyst. The money goes to capacity. Snorkel plans to expand its agentic data factory so production can rise with demand, hire researchers and engineers, accelerate investment in vertical and enterprise AI, and extend the work into new domains and modalities. Enterprise and government operations will grow with the factory. The company also plans to support third-party evaluations of AI models.
"The teams pushing the frontier want a research data partner who pioneers the science of data development," Ratner said. "That's what Snorkel was built to be: the frontier lab for agentic data, combining human excellence with over a decade of research and technology." Growth remains the priority. The company expects to reach profitability by the end of 2026.
In a blog post published with the raise, Ratner reported an annualized revenue run-rate of $375 million. A year earlier the figure had been roughly $20 million. Roughly 157 people worked at the company, which put revenue per employee at about $2.3 million. It had grown 18 times over.
How it spread
Moving into data services with all those specialists makes sense but it feels like a big pivot from their original software focus.
Snorkel hitting $350M ARR this fast shows how desperate the big AI labs are for reliable training data right now.