Snorkel AI is a San Francisco company spun out of the Stanford AI Lab in 2019 and known for the Snorkel weak-supervision project. It now calls itself 'the frontier AI data lab'. In 2025 it moved from selling Snorkel Flow software to an expert Data-as-a-Service model, and it now supplies expert agentic tasks, RL environments (computer-use, terminal and simulated-enterprise), rubrics and evals to AI labs and enterprises. It combines expert contributors with synthetic generation and agent-based QC, maintains Terminal-Bench, and co-authors or hosts agent benchmarks such as OSWorld 2.0, Agents' Last Exam and Senior SWE-bench through a $3M Open Benchmarks Grants program. It raised a $350M Series E at $3.5B in September 2026 and says its ARR passed $375M.
Snorkel's blog covered MedPAIR, built by researchers from MIT, Oxford, Google Research and Cornell (lead: Yuexing Hao). It collects sentence-level relevance annotations from 52 physician trainees and 7 board-certified physicians on 933 medical exam questions and compares them with six LLMs. Humans and models agree on roughly 50-60% of sentences. Snorkel funded the work through its Open Benchmarks Grants and helped coordinate physician recruitment.
Why it matters: Shows Snorkel funding healthcare evaluation research. Snorkel did not author the dataset.
An explainer by Aryan Kargwal and Jonathan Schlosser on how an environment's state, observations, actions, transitions, rewards and reset shape what agents learn. It discusses public environments including Terminal-Bench 2.0, CUA-Gym, OSWorld 2.0 and SCUBA.
Why it matters: Sets out how Snorkel approaches environment and reward design. It announces no new release.
Co-led by Insight Partners and S32, with significant participation from Addition. New investors are March Capital, Blumberg Capital, Allegis Capital, Frontline, Standard and Third Point Ventures. Returning investors are Greylock, Lightspeed, GV, Factory, Prosperity7, Walden Catalyst and Wells Fargo. Snorkel says ARR passed $375M, up more than 18x in under a year since it launched expert Data-as-a-Service in September 2025. The money will expand its 'frontier AI data factory' for labs and enterprises, with investors naming healthcare, law and software engineering.
Why it matters: Puts Snorkel among the scaled incumbents in frontier-lab data and environments. TechCrunch names Mercor, Handshake and micro1 as peers. The ARR figure is self-reported.
TechCrunch (Marina Temkin) reports a $375M annualized run rate, up eighteenfold over 12 months, and a prior valuation of $1.3B at the Series D 17 months earlier. It describes the move from data-labeling software to data-as-a-service, which combines synthetic generation with subject-matter experts. It notes that expert payments sit in cost of goods sold rather than in headline revenue, unlike pure labor marketplaces. Reuters, reprinted on Snorkel's press page, puts the run rate above $350M, up from about $20M a year earlier.
Why it matters: Independent press backing for the round and the revenue scale. The run-rate figures differ slightly between outlets ($350M+ vs $375M).
Alex Ratner writes that Snorkel partners with 'leading frontier labs, hyperscalers, neolabs, vertical AI leaders, enterprises, and U.S. government agencies'. He says that when building coding-agent environments and datasets, its internal Agentic Data Platform runs hundreds of specialized QC agents alongside human experts. The claimed results are 50%+ better QC efficiency than human-only review and 15+ accuracy points over a human paired with an off-the-shelf LLM.
Why it matters: Shows how Snorkel mixes humans and AI in production. All performance figures are self-reported, and no customers are named.
Snorkel describes Terminal-Bench+ as an extension of Terminal-Bench into 'a training-scale data series organized as a progressive curriculum'. The company says it contains more than 50,000 expert-authored tasks across 12 task types and 9 languages, with codebases of 200+ files. In an analysis of 740 trajectories from GPT-5.5 and Claude Opus 4.8, skipping verification caused 43% of failures. Samples are available on request, and licensing terms are not stated.
Why it matters: A commercial coding and terminal environment/data product aimed at RL training, directly relevant to buyers of coding environments. Task counts are self-reported.
Fortune's article 'OpenAI quietly boosts some of Astra's evaluation metrics, and continues to change others post-launch' quoted Vincent Sunn Chen. He said benchmark results depend on the exact model checkpoint, compute budget, harness and evaluation configuration, and called for clearer disclosure of changes to them.
Why it matters: Shows Snorkel's public profile on evaluation and benchmarking.
Snorkel's post on OSWorld 2.0 (108 long-horizon tasks averaging 300+ steps) reports the paper baseline: Claude Opus 4.8 completes 20.6% outright with 54.8% partial credit. On the Snorkel-hosted leaderboard, Claude Fable 5.1 reaches about 45% binary completion and above 60% partial credit. Snorkel supported the benchmark through its Open Benchmarks Grants and hosts the leaderboard.
Why it matters: Snorkel is a co-author and leaderboard host for a major computer-use agent benchmark.
Snorkel says it maintains Terminal-Bench, which it calls 'one of the most widely reported benchmarks on frontier model cards', as a continuously maintained benchmark. Version 4.0 pins base and built images, raises timeouts and CPU/memory budgets, fixes or removes tasks, and calibrates the whole benchmark. Post-launch QA found that some failures came from resource budgets rather than capability gaps, and a solvability study found well-formed tasks that could not be solved. The Snorkel-hosted 4.0 leaderboard lists 66 tasks.
Why it matters: Snorkel maintains a coding and terminal agent benchmark that frontier labs report on, which matters to buyers of coding environments.
Snorkel describes environments for insurance, finance, manufacturing, legal and sales/marketing. The insurance underwriting environment has 50+ database tables, 100 documents, 47 workflow tools and 7 user personas. RL-training Qwen3-30B on it raised BFCL Multi-Turn from 33.3 to 41.4 (+8.1), τ³-bench overall from 24.6 to 30.4 (+5.8) and τ³-bench airline from 32.0 to 46.0 (+14.0).
Why it matters: Public evidence that Snorkel's environments are used for RL training, not only evaluation, and that gains transfer out of domain. The results are self-reported.
The projects are Frontier-Bench (with Laude Institute and Harbor), Agents' Last Exam (UC Berkeley RDI), OSWorld 2.0 (XLANG Lab), Continual Learning Bench (UC Berkeley SkyLab and UW-Madison), SlopCode Bench (UW-Madison), Terminal-Bench 2.1 (Stanford, Harbor and Laude Institute), Terminal-Bench Science (in development) and Senior SWE-Bench (Princeton and UW-Madison). Supporting partners include Hugging Face, Prime Intellect, Together AI, Factory, Harbor and PyTorch. The program launched in February 2026 and takes applications on a rolling basis.
Why it matters: Puts Snorkel at the center of open agent-benchmark development, which is also how it reaches frontier labs.
An open, Harbor-native benchmark of underspecified, senior-level design-and-build and investigate-and-fix tasks. It has 100 tasks (50 public, 50 private) from 12 open-source repositories. Henry Kiss Ehrenberg led the project and Vincent Sunn Chen co-authored it. The release says top frontier models fail to reach senior-level correctness and taste more than 75% of the time. The live site now says more than 65%. The dataset is on GitHub at snorkel-ai/senior-swe-bench-v2026.06.
Why it matters: A new coding-agent benchmark co-built by Snorkel and published in Harbor format.
Authors include Justin Bauer, Thomas Walshe, Derek Pham, Harit Vishwakarma, Armin Parchami, Frederic Sala and Paroma Varma. They study RL with verifiable rewards on small models using procedurally generated counting, graph and spatial reasoning tasks. Mixed-complexity training gives up to 5x sample efficiency over training on easy tasks. Snorkel's research page lists the paper as MLSys 2026; arXiv shows no venue.
Why it matters: Snorkel's own RLVR research on data difficulty and curriculum, which supports its pitch for curriculum-style environments.
| Date | Round | Amount | Led by | |
|---|---|---|---|---|
| 22 Sep 2026 | Series E | $350M valuation $3.5B | Insight Partners, S32 | source ↗ |
| 6 Aug 2025 | Strategic investment | undisclosed | Accenture | source ↗ |
| undisclosed | ||||
| 29 May 2025 | Series D | $100M valuation $1.3B | Addition | source ↗ |
| 9 Aug 2021 | Series C | $85M valuation $1B | BlackRock, Addition | source ↗ |
| 7 Apr 2021 | Series B | $35M | Lightspeed Venture Partners | source ↗ |
| Jul 2020 | Seed + Series A | $15M | – | source ↗ |
| $15M (seed + Series A combined) | ||||