RL List.com
UPDATED 2026.10.05
rl-list.com · Vendors · micro1

micro1

Incumbent
medium confidence
www.micro1.ai ↗ · Status: active confirmed · Founded 2022

micro1 is a San Francisco AI training-data company founded in 2022 by Ali Ansari. It began as an AI recruiter (Zara) and now uses that system to vet expert contractors. It sells expert human data, frontier evaluations and RL environments to AI labs under the Realm brand, an enterprise agent-evaluation layer called Cortex, and expert-demonstrated robotics data, all run on its flow data platform. To make its 'RL gyms' realistic it pays companies for operational data, and it made a $12.5M late bid for Spirit Airlines' records. TechCrunch reported a $500M gross run rate in August 2026, and Forbes reported a raise of more than $100M at a $4B valuation in September 2026, with frontier labs taking part.

Key facts
Headquarters
San Francisco, CA, USAreported cite
Headcount band
51-200reported cite
Total raised
$138Mestimated cite
Last round
$100M+ at $4B valuation, 2026-09 (stage not disclosed)reported cite
SOC 2
unknown cite
What they sell
human data + environmentsconfirmed cite
Open source
partial (LongExtractionBench code and sample data; RedlineBench dataset with Crosby; Realm, Cortex and flow are proprietary)confirmed cite
Deployment
managed service (expert data, evals and RL environments delivered to labs and enterprises)estimated cite

What's new

13 sourced updates · last 6 months

Founding team

Ali Ansari
Ali Ansari
Founder & CEO
Born in Iran and raised in Los Angeles. Studied computer science at UC Berkeley, and his micro1 bio adds Stanford computer science with a focus on reinforcement learning (he told the Stanford Daily in 2025 that he was a master's student). Before micro1 he ran a textbook-resale business, a math-tutoring platform and a software agency. Forbes calls him a solo founder.
Andrew Maas
VP of AI (non-founder)
VP of AI at micro1, co-author of the September 2026 flow-transform PII paper and the robotics safety research. Prior background is not stated on micro1 pages.
Yongchao Zhou
Director of Research (non-founder)
Director of Research at micro1, per the September 2026 research byline. Prior background is not stated on micro1 pages.
Kathy Wang
Kathy Wang
Chief Information Security Officer (non-founder)
micro1's CISO. Her 'Why I joined micro1' post says the expert data passing through micro1 is among the most sensitive in the AI supply chain. Prior roles are not stated on the page.

What practitioners should know

For researchers

  • The Realm benchmark pages (legal, financial, tax, pathology v1 and v2) document the task design: sandboxed agents with shell, file I/O and web search; 35-60 expert rubric items per legal task, and 1,071 criteria in total for pathology v2; and an LLM judge giving a reward from 0 to 1. Most of these pages do not link a public dataset or paper.
  • Open releases are limited: LongExtractionBench (github.com/micro1-research/longextract-bench, MIT; Hugging Face micro1-inc/longextract-bench-50), RedlineBench with Crosby (Hugging Face crosbylegal/RedlineBench, CC-BY-4.0) and a small Hugging Face dataset, micro1-inc/Prospera_Benchmark. No open environments or model weights were found.
  • Papers with arXiv versions: 'No Last Mile' (2603.00932) and 'The Benchmark Ceiling' (2607.01254). The flow-transform PII paper and the robotics-safety study (both September 2026) appear only on micro1.ai. In a 2025 Stanford Daily interview, Ansari said Stanford's Stefano Ermon leads micro1's research lab. No 2026 micro1 page confirms this.

For program managers

  • Scale: TechCrunch, citing an anonymous source, reports a $500M gross run rate (August 2026) and estimates net at $150M-$200M. Earlier figures: about $7M ARR at the start of 2025, $50M in September 2025 and $100M+ in December 2025.
  • Frontier-lab ties: Forbes (September 2026) names unnamed frontier labs, Microsoft, Amazon and 1X as customers, and reports that two frontier labs and two xAI co-founders joined the $100M+ round at $4B. TechCrunch (December 2025) described Microsoft as one of the AI labs micro1 works with. micro1's own site names no lab.
  • Engagement is a managed service: Realm for labs, Cortex for enterprises and Robotics for embodied AI, all run on the flow platform. The main named public reference is a self-published Box case study (enterprise agent eval data), which has no Box quote.
  • micro1 makes its environments realistic by buying enterprise data: data partnerships at $100k to $1M+, and a $12.5M late bid for Spirit Airlines' records that was unresolved in coverage through late September 2026. Check data provenance and consent terms.

For data ops

  • No public trust center and no SOC 2 or ISO 27001 claim was found on the micro1.ai home, Realm, government or data-partnerships pages, and /security and trust.micro1.ai were not found. micro1 does have a named CISO, Kathy Wang. Ask for attestations directly.
  • Privacy controls stated on the data-partnerships page: agreed data scope, PII removal and anonymization, restricted access, isolated processing pipelines, and deletion on request. flow-transform 1.0 is the de-identification tool.
  • No public API docs or delivery-format specs were found. Government standing: CDAO Tradewinds 'Awardable' status, and micro1 says it was selected to support the DOE Genesis Mission (September 2026).

Benchmarks & research

12 published
Corpus-level PII transformation, scored with a Transformation Quality Index (TQI) and detection F1.
TQI 88.9 with agentic review vs 74.9 for NVIDIA NeMo Anonymizer. Detection F1 0.829 vs 0.678. Self-reported.
Fault-injection tests of Claude Opus 5 and Sonnet 5 controlling robots on lab tasks.
Both models failed both tasks, scoring 6/15 to 9/15
Agents read de-identified anatomic-pathology reports and write structured summaries. About 80% of tasks use a single report and 20% combine several. Expert-authored rubrics hold 1,071 criteria (772 positive, 299 penalty), and an LLM judge (GPT-5.4 mini) scores each one.
Claude Fable 5.1 leads the main scoreboard at 78.0%, ahead of Muse Spark 1.3 (75.8%) and Claude Opus 5 (75.2%). A separate summary table lists Claude Opus 5.5 at 77.0%. Extraction criteria are met 93% of the time, interpretation criteria 73%.
The Benchmark Ceilingpaper · 13 Jul 2026
Esposito (micro1) and Zhang (Harvard) on how benchmarks depreciate and why expert evaluation labor is scarce.
LongExtractionBenchbenchmark · 30 Jun 2026
Seven extraction systems compared on 225 dense documents averaging 358 pages. Grading code is MIT-licensed on GitHub, with a 50-document sample on Hugging Face.
Reducto Deep Extract completed 225/225 at 99.6% precision and recall. Claude Opus 4.8 and Gemini 3.1 Pro completed 116 and 112 of 225.
50 tasks (41 single-report, 9 multi-report) across eight clinical domains. Agents get shell, file I/O, an editor and web search, and an LLM judge grades them on expert rubrics.
Claude Opus 4.8 82.6%, GPT-5.5 76.3%, Gemini 3.5 Flash 75.7%
Realm: Financial reasoning benchmarkbenchmark · 15 May 2026
103 expert-authored finance tasks with 144 attachments, grounded in real work products. 927 rollouts across GPT-5.5, Claude Opus 4.7 and Gemini 3.1 Pro.
GPT-5.5 leads on mean reward. No model clears 50% on judgment-call tasks, and VC tasks average 22.8%.
Realm: Legal reasoning benchmarkbenchmark · 10 Apr 2026
Long-horizon litigation, transactional and compliance tasks in a sandbox. Each task has 35-60 rubric items decomposed along IRAC lines, and an LLM judge gives a reward from 0 to 1.
Claude Opus 4.7 and GPT-5.5 are statistically tied (0.007 apart). Gemini 3.1 Pro scores about two-thirds of either.
Ansari, Esposito, Fitoussy and Zhang model structured human data work (evaluation, auditing, QA) as a lasting input to AI production rather than a temporary one.
Predicts a long-run structured-labor share of 5-7%
Agents prepare a complete federal (Form 1040) return with schedules from source documents, scored against expert-authored criteria.
GPT-5.4 leads with mean reward 0.368 and pass@3 of 28%. 44% of criteria go unsolved by every model.
Multi-turn SaaS MSA negotiation and redlining benchmark built with Crosby Legal, using attorney-authored rubrics across five dimensions. The dataset is on Hugging Face (crosbylegal/RedlineBench, CC-BY-4.0).
Turn-weighted scores: GPT-5.5 50.5%, Claude Fable 5 47.3%, Gemini 3.5 Flash 45.1%, Claude Opus 4.8 44.4%
CortexRetrievalBenchbenchmark · 2026
Financial-research retrieval benchmark: 140 tasks over 155 documents from 29 companies. Tests whether retrievers rank the authoritative source first.
Success@1 ranges from 42.1% (BM25) to 69.3% (OpenAI dense retriever). Authority Success@10 ranges from 92.1% to 99.3%.

Products

Realm ↗
Sold as 'Frontier evaluations and RL environments for AI Labs': expert-authored, sandboxed agent tasks with rubric-based rewards across 100+ fields (finance, medical, legal, coding, STEM, VLM, audio), delivered as a managed service. Forbes describes these as RL gyms built with real enterprise data.
Cortex ↗
Evaluation and improvement layer for enterprise agents in production: eval design, failure diagnosis by domain experts, targeted training data and reliability monitoring.
Robotics data ↗
Expert-demonstrated POV and third-person capture, 2D/3D annotation and QC for embodied-AI labs.
flow ↗
Data platform that pairs expert judgment with model-assisted task generation (flow gen), QC (flow qc), grading (flow grader) and orchestration. Variants include flow-transform for PII and Flow for mission data (defense EO/IR).launched 22 Sep 2026
Data partnerships ↗
micro1 pays companies for operational data such as SOPs, knowledge bases, workflow and decision data, and human feedback. Tiers are $100k+, $500k+ and $1M+. Data goes through PII removal, restricted access and isolated pipelines, and is used for Realm evaluations and RL environments.
Zara (AI interviewer) and expert network ↗
AI recruiter that vets the expert contractors. The government page cites '130,000+ deeply vetted candidates across 100+ domains and 60+ languages'.
Government ↗
Agent buildouts and human data for government buyers. micro1 holds 'Awardable' status on the CDAO Tradewinds Solutions Marketplace.

Scale & velocity

Current headcount
Full-time staff not disclosed. LinkedIn lists a 51-200 band, but about 10,400 people list micro1 on LinkedIn, likely including contractors.reported cite
Headcount growth
unknown
Open roles
unknown cite
Other locations
unknown
Distributed / remote
unknown

Research depth

Has researchers
yesconfirmed cite
Researcher count
unknown
Backgrounds
Andrew Maas, VP of AI, Yongchao Zhou, Director of Research, Rumi Elias Calles, Director of Research Engineering, Arian Sadeghi, VP Robotics, Mark Esposito, Chief Economist (co-authors with Harvard's Liu Zhang), Stanford's Stefano Ermon, said by Ansari to lead micro1's research lab (Stanford Daily, 2025-10; not confirmed on micro1 pages)reported cite

Capital

Total raised
$138M+ (est. sum: $3.3M pre-seed/seed 2023 + $35M Series A 2025-09 + $100M+ round 2026-09)estimated cite
Last round
$100M+ at $4B valuation, 2026-09 (stage not disclosed)reported cite
Investors
01 Advisors (Dick Costolo, Adam Bain; led Series A; Bain on board), Two unnamed frontier AI labs (2026 round, reported), Two unnamed xAI co-founders (2026 round, reported), Dream Ventures (led 2023 round), Jason Calacanis (2023), Joshua Browder (2023; board member), Cory Levy (2023)reported cite
Valuation
$4B (2026-09, reported)reported cite
Revenue signals
$500M gross annual run rate (2026-08), with net estimated at $150M-$200M (TechCrunch, citing a person familiar with the company). Earlier: about $7M ARR at the start of 2025, $50M in 2025-09, $100M+ in 2025-12.reported cite
DateRoundAmountLed by
22 Sep 2026undisclosed$100M
valuation $4B
–source ↗
$100M+ ('more than US$100 million')
12 Sep 2025Series A$35M
valuation $500M
01 Advisors (Dick Costolo, Adam Bain)source ↗
24 Oct 2023Pre-seed / seed$3.3M
valuation $30M
Dream Venturessource ↗
$3.3M total ($2M pre-seed plus a $1.3M oversubscribed round)

Security & compliance

SOC 2
unknown cite
Other certifications
CDAO Tradewinds Solutions Marketplace 'Awardable' status (procurement status, not a security certification)confirmed cite
Security page
unknown

Product

What they sell
human data + environmentsconfirmed cite
Open source
partial (LongExtractionBench code and sample data; RedlineBench dataset with Crosby; Realm, Cortex and flow are proprietary)confirmed cite
License
mixed: longextract-bench MIT; RedlineBench dataset CC-BY-4.0 (Crosby); commercial products proprietaryconfirmed cite
Deployment model
managed service (expert data, evals and RL environments delivered to labs and enterprises)estimated cite
Maturity
GAreported cite
Notable customers
⚑Microsoft verified ⚑Amazon self-claimed 1X self-claimed Box self-claimed cite

Buyer analysis

Best fit: Labs and enterprises that want a large managed vendor for expert-authored, rubric-graded agent tasks and evals in professional domains (legal, finance, tax, healthcare), for RL environments seeded with real enterprise data, or for robotics demonstration data.

How we verified this

2026-10-05 initial profile: I opened and checked the core claims against their sources: - **Forbes Australia syndication:** more than $100M at $4B; two frontier labs and two xAI co-founders in the round; customers including frontier labs, Microsoft, Amazon and 1X; micro1 declined to comment. - **TechCrunch:** the 2025 Series A of $35M at $500M led by 01 Advisors; the December 2025 $100M ARR article describing Microsoft as an AI lab customer; the August 2026 $500M gross run rate from an anonymous source, with a $150M-$200M net estimate. - **Pulse 2.0:** the 2023 $3.3M round at a $30M valuation. - **micro1 pages:** flow, DOE Genesis, the defense EO/IR post, the PII paper, The Benchmark Ceiling, No Last Mile, all Realm benchmark pages, data partnerships and the government page. - **GitHub and Hugging Face:** the micro1-research and micro1-inc orgs, and the Crosby RedlineBench dataset. - **Spirit coverage:** TNW, TIME and Anadolu. Both founder headshots were downloaded and viewed. Each shows one person and comes from micro1's own pages. Corrections: - Pathology v2 headline: Opus 5.5 at 77.0% appears only in a secondary table. - TechCrunch's retention wording contradicts its own net figure, so the item now flags it. - The Spirit bid timeline and court status were reframed. - Precise dates were added for several items. Items the researcher missed and I added: the September 2026 robotics safety research (with VP Robotics Arian Sadeghi), the No Last Mile arXiv paper, and the RedlineBench CC-BY-4.0 license. Remaining gaps: - No SOC 2 or trust center was found. - No frontier lab is named as a customer by any primary source. - The 2026 round and the revenue figures rest on anonymous-source reporting. - No court ruling on Spirit was found. - WebSearch was unavailable (budget exhausted), so discovery used Google News RSS. Forbes.com itself was not opened; its Australian syndication was used instead.

Related vendors

Sources

  1. www.micro1.ai/ · 2026-10-05, Homepage
  2. www.micro1.ai/realm · 2026-10-05, Realm: frontier evals and RL environments for AI labs. No compliance badges, no named labs.
  3. www.micro1.ai/cortex · 2026-10-05, Cortex enterprise agent eval layer
  4. www.micro1.ai/robotics · 2026-10-05, Robotics data offering
  5. www.micro1.ai/data-partnerships · 2026-10-05, Pays $100k+, $500k+ or $1M+ for operational data. Privacy controls.
  6. www.micro1.ai/government · 2026-10-05, CDAO Tradewinds Awardable. 130,000+ vetted candidates. No SOC 2 or FedRAMP.
  7. www.micro1.ai/research · 2026-10-05, Research index, including No Last Mile (2026-03) and the robotics safety study (2026-09)
  8. www.micro1.ai/research/no-last-mile · 2026-10-05, 2026-03-01, arXiv 2603.00932
  9. www.micro1.ai/research/ai-models-now-introduce-safety-risks-in-the-physi · 2026-10-05, 2026-09-18. Arian Sadeghi, VP Robotics.
  10. www.micro1.ai/newsroom · 2026-10-05, Press list, including Forbes AI 50 Brink (2026-04-16)
  11. micro1.ai/blog/introducing-flow · 2026-10-05, flow platform, 2026-09-22
  12. micro1.ai/blog/micro1-supporting-the-doe-genesis · 2026-10-05, DOE Genesis selection, 2026-09-01
  13. micro1.ai/blog/data-for-our-defense-customers · 2026-10-05, Flow for mission data, 2026-08-02
  14. www.micro1.ai/research/pii-transformation-for-enterprise-datasets · 2026-10-05, flow-transform 1.0, 2026-09-24
  15. www.micro1.ai/research/the-benchmark-ceiling · 2026-10-05, arXiv 2607.01254, 2026-07-13
  16. www.micro1.ai/benchmark/realm-legal · 2026-10-05, Legal benchmark, 2026-04-10
  17. www.micro1.ai/benchmark/realm-financial · 2026-10-05, Financial benchmark, 2026-05-15
  18. www.micro1.ai/benchmark/realm-tax · 2026-10-05, Tax benchmark
  19. www.micro1.ai/benchmark/realm-pathology-report · 2026-10-05, Pathology v1 (June 22)
  20. www.micro1.ai/benchmark/realm-pathology-report-v2 · 2026-10-05, Pathology v2 (reauthored 17-20 September 2026)
  21. www.micro1.ai/benchmark/crosby-micro1-redlinebench · 2026-10-05, RedlineBench with Crosby
  22. huggingface.co/datasets/crosbylegal/RedlineBench · 2026-10-05, CC-BY-4.0 dataset owned by crosbylegal
  23. www.micro1.ai/benchmark/cortex-retrieval-bench · 2026-10-05, CortexRetrievalBench
  24. www.micro1.ai/benchmark/long-extraction · 2026-10-05, LongExtractionBench
  25. github.com/micro1-research · 2026-10-05, One public repo, longextract-bench (MIT), updated 2026-06-30
  26. huggingface.co/micro1-inc · 2026-10-05, Two datasets, no models
  27. www.micro1.ai/ali-ansari · 2026-10-05, Ansari bio. Headshot checked: HTTP 200 image/webp, single-person portrait.
  28. www.micro1.ai/team/why-i-joined-micro1-2 · 2026-10-05, Kathy Wang, CISO. Author headshot checked: HTTP 200 image/png, single-person portrait.
  29. www.micro1.ai/case-study/box · 2026-10-05, Box case study (self-published, no Box quote)
  30. www.micro1.ai/series-a · 2026-10-05, Series A: $35M at $500M post-money, 2025-09-12
  31. www.forbes.com.au/news/uncategorized/this-25-year-old-raised-us100-milli · 2026-10-05, Forbes (Anna Tong) syndication: more than $100M at $4B, per two people familiar with the deal
  32. techcrunch.com/2026/08/20/ai-data-startup-micro1-reaches-500m-gross-run- · 2026-10-05, $500M gross run rate, $150M-$200M net, anonymous source
  33. techcrunch.com/2025/12/04/micro1-a-scale-ai-competitor-touts-crossing-10 · 2026-10-05, $100M ARR. 'works with leading AI labs, including Microsoft'. RL work.
  34. techcrunch.com/2025/09/12/micro1-a-competitor-to-scale-ai-raises-funds-a · 2026-10-05, $35M Series A led by 01 Advisors at $500M. Bain and Browder on the board.
  35. finance.yahoo.com/news/scale-ai-rival-micro1-hits-134611177.html · 2026-10-05, Founded 2022. $50M ARR. Expanding into environments.
  36. pulse2.com/micro1-ai-recruiting-company-raises-3-3-million/ · 2026-10-05, 2023-10-24: $3.3M at $30M valuation
  37. e.vnexpress.net/news/tech/personalities/from-50-ebay-flip-to-2-5b-ai-sta · 2026-10-05, 2026-02-21: $2.5B valuation per Forbes estimates. Ansari background.
  38. stanforddaily.com/2025/10/16/micro1-founder-ali-ansari-on-ai-and-human-i · 2026-10-05, Ansari says Stefano Ermon leads micro1's research lab
  39. thenextweb.com/news/micro1-12-5m-counterbid-spirit-airlines-records-goog · 2026-10-05, Spirit counter-bid details, 2026-09-04
  40. www.aa.com.tr/en/science-technology/ai-firms-set-sights-on-data-from-ban · 2026-10-05, 2026-09-30: approval hearing still upcoming
  41. time.com/article/2026/08/25/google-spirit-airlines-ai-data-RL/ · 2026-10-05, Mercor $7.5M. micro1's bid came after the auction closed.
  42. www.linkedin.com/company/micro1/ · 2026-10-05, 51-200 band, SF Bay Area, founded 2022, about 10,400 associated members
Last updated 2026-10-05 · Every quantitative field carries a source and a confidence tag. Fields we could not source publicly are marked unknown, never estimated. See the methodology.