micro1 is a San Francisco AI training-data company founded in 2022 by Ali Ansari. It began as an AI recruiter (Zara) and now uses that system to vet expert contractors. It sells expert human data, frontier evaluations and RL environments to AI labs under the Realm brand, an enterprise agent-evaluation layer called Cortex, and expert-demonstrated robotics data, all run on its flow data platform. To make its 'RL gyms' realistic it pays companies for operational data, and it made a $12.5M late bid for Spirit Airlines' records. TechCrunch reported a $500M gross run rate in August 2026, and Forbes reported a raise of more than $100M at a $4B valuation in September 2026, with frontier labs taking part.
The authors are Aruj Mahajan (AI Lead Engineer), Andrew Maas (VP of AI), Yongchao Zhou (Director of Research) and Rajeev Joshi. flow-transform 1.0 replaces real identities with consistent synthetic ones across an enterprise corpus. It scores a Transformation Quality Index of 88.9 with agentic review and 84.1 without, against 74.9 for NVIDIA's NeMo Anonymizer baseline. Detection F1 is 0.829 with agentic review, 0.758 without, and 0.678 for NeMo.
Why it matters: This is the de-identification layer behind micro1's purchases of enterprise operational data for RL environments. All results are self-reported.
Forbes (Anna Tong), citing two people familiar with the deal, reports a round of more than $100M at a $4B valuation, up from $500M in September 2025. Participants include two frontier labs and two xAI co-founders. No lead is named, and micro1 declined to comment. Forbes says customers include frontier AI labs, Microsoft, Amazon and robotics companies such as 1X. It describes micro1's 'reinforcement learning gyms' as simulated environments, built with real-world enterprise data, where AI agents can practise navigating workplaces.
Why it matters: Frontier labs investing in micro1 as well as buying from it is a strong demand signal. The round comes from anonymous sources, not a company announcement.
Rumi Elias Calles (Director of Research Engineering) and Tyler Houchin describe four flow components: flow gen (drafts and task variations), flow qc (quality inspection, such as catching prompt/rubric contradictions), flow grader (grading outputs against criteria or ground truth) and flow orchestration (a shared catalog of model capabilities). micro1 says flow underpins Realm (lab evals and RL environments), Cortex (enterprise agents) and Robotics.
Why it matters: Shows how micro1 uses model-assisted generation, QC and grading to scale expert-authored RL tasks and rubrics.
Arian Sadeghi (VP, Robotics), Andrew Maas, Mitali Potnis and Ali Ansari injected faults mid-task in two robotic workflows: toxic liquid handling, and loading a sample tube into an analyzer. They tested Claude Opus 5 and Sonnet 5. Both failed. Opus 5 detected anomalies sooner but still caused damage in the manipulation task. Scores ranged from 6/15 to 9/15.
Why it matters: Shows micro1's embodied-AI evaluation work alongside its robotics data offering. The results are self-published.
micro1 offered $12.5M for Spirit's internal records after Google's $10M bid had won the auction. TIME (2026-08-25) reported that micro1's bid came after the auction closed and that Mercor had bid $7.5M. Per TNW, the records include about 500M Microsoft Teams items, 100M emails and roughly 16M customer chat sessions. micro1 proposed an ombudsman chosen by Spirit's advisers, U.S.-based storage, and exclusion of disciplinary, investigatory and union-related records. No court ruling was found in coverage through 2026-09-30 (Anadolu).
Why it matters: Shows micro1 buying real enterprise data to make its RL environments realistic, and the privacy scrutiny that comes with it. The outcome is unresolved in the sources found.
micro1 announced that it was selected to contribute to the U.S. Department of Energy's Genesis Mission. Its stated role covers making complex scientific and technical data usable for AI, building evaluation workflows, and bringing in expert knowledge. No contract value was disclosed.
Why it matters: A federal science engagement, alongside micro1's CDAO Tradewinds 'Awardable' status. This is micro1's own announcement and has no independent confirmation.
Citing a person familiar with the company, TechCrunch (Marina Temkin) reports that micro1's gross annual run rate grew from $100M to $500M over eight months. It puts net run rate at $150M-$200M. The article says micro1 'retains roughly 60% to 70%' of gross, which does not match its own $150M-$200M net figure, since that range implies about 30-40% retained. Products cited are RL gyms, a robotics pretraining dataset, synthetic data generation, and off-the-shelf datasets at 80-90% gross margins. TechCrunch also says micro1 'may have recently raised another round at a significantly higher valuation'.
Why it matters: A scale signal. The headline '$500M' is gross bookings, not net revenue, and it comes from an anonymous source.
micro1 applies its flow pipeline to aerial electro-optical/infrared footage, with human verification layers on top. It cites a threefold reduction in human correction load at constant quality, a figure carried over from its published robotics work. No agencies are named.
Why it matters: Expands micro1 into defense data. No security accreditations are stated.
Mark Esposito (micro1) and Liu Zhang (Harvard) model how benchmarks lose value as they saturate, and document a scarcity premium for expert evaluation labor. They use micro1 platform data on more than 1,000 credentialed professionals (arXiv 2607.01254).
Why it matters: micro1's argument that expert-built evaluations stay scarce and valuable as benchmarks saturate.
Seven extraction systems compared on 225 documents averaging 358 pages. Grading code is at github.com/micro1-research/longextract-bench (MIT), and a 50-document sample is at huggingface.co/datasets/micro1-inc/longextract-bench-50. The full corpus is not released for licensing reasons. The date is when the repo and dataset were last updated, since the benchmark page shows none.
Why it matters: The only public code repository from micro1 that we found.
103 expert-authored finance tasks with 144 attachments and 927 rollouts across GPT-5.5, Claude Opus 4.7 and Gemini 3.1 Pro. GPT-5.5 leads on mean reward. No model clears 50% on tasks that need a judgment call, and venture-capital tasks averaged 22.8%.
Why it matters: A public sample of the expert-rubric agent tasks micro1 sells to labs under Realm.
micro1's newsroom lists its inclusion on the Forbes AI 50 Brink list, dated 2026-04-16.
Why it matters: Recognition signal only. Sourced from micro1's newsroom.
Long-horizon litigation, transactional and compliance tasks. Each has 35-60 rubric items, split as issue 4%, rule 33%, application 48%, conclusion 9% and other 6%, and an LLM judge gives a reward from 0 to 1. Claude Opus 4.7 and GPT-5.5 are statistically tied (0.007 apart), and Gemini 3.1 Pro scores about two-thirds of either.
Why it matters: A public example of how micro1 structures its legal RL and eval tasks.
| Date | Round | Amount | Led by | |
|---|---|---|---|---|
| 22 Sep 2026 | undisclosed | $100M valuation $4B | – | source ↗ |
| $100M+ ('more than US$100 million') | ||||
| 12 Sep 2025 | Series A | $35M valuation $500M | 01 Advisors (Dick Costolo, Adam Bain) | source ↗ |
| 24 Oct 2023 | Pre-seed / seed | $3.3M valuation $30M | Dream Ventures | source ↗ |
| $3.3M total ($2M pre-seed plus a $1.3M oversubscribed round) | ||||