#24
Idler
Commercial
medium confidence
idler.ai ↗ · Status: active
confirmed · Founded 2025
Idler (YC S25, founded 2025, San Francisco) calls itself a frontier data research lab that builds evals and RL environments from real production work. Its public benchmarks are ShelfLife, a digital twin of a live multi-brand e-commerce company; ShelfLife E-Sim, a stateful retail-operations simulator; and CorpLaw, built from a law firm's anonymized data. It also offers private, on-request collections in long-horizon SWE, cybersecurity, recursive self-improvement and terminal agents. The roughly 12-person team is led by Dark Forest/0xPARC co-founder Ivan Chub and says leading frontier labs use its work, without naming any. A $9M Paradigm-led seed was reported by an aggregator in August 2026 but is not confirmed by the company.
Key facts
Headquarters
San Francisco, CA, USA
confirmed citeLast round
Seed, $9M, 2026-08-21, led by Paradigm with Y Combinator and Long Journey
reported citeWhat they sell
environments + evals
confirmed citeDeployment
Bespoke/custom environments and private datasets delivered to labs on request. Public leaderboards, docs and samples are on idler.ai. The delivery mechanism is not publicly documented.
confirmed cite
What's new
4 sourced updates · last 6 months- 26 Aug 2026benchmark
ShelfLife E-Sim leaderboard published: long-horizon retail-operations simulator
Idler published ShelfLife E-Sim, a stateful simulator of retail operations grounded in 'nearly a decade of real customer orders, purchase orders, and inventory snapshots from a live retailer'. Agents take control of the business on a fixed date and run it one day at a time. On the measurement of 2026-08-26, Grok 4.6 had the top mean reward (65.0 ± 1.7 pts), followed by GPT-5.6 Sol (54.1), Gemini 3.7 Flash (51.7), Fable 5 (51.0), Opus 5 (39.5) and Kimi K3 (37.6).
Why it matters: This is a long-horizon, turn-based business simulation, and its model rankings differ sharply from the static ShelfLife tasks: Opus 5 leads ShelfLife but places fifth here.
- 21 Aug 2026funding
Idler raises a $9M seed from Paradigm to build RL environments (reported)
The SaaS News reports a $9M seed led by Paradigm, with Y Combinator, Long Journey and angel investors. The money will go to growing the ML research team, scaling compute, and building more complex RL environments.
Why it matters: This is the company's first reported priced round, but it rests on a single aggregator headline.
- 18 Aug 2026benchmark
CorpLaw: a corporate-law benchmark built from a real law firm's anonymized data
CorpLaw places agents in transaction workspaces with documents, emails and client instructions. Grading criteria are 'grounded in the work the firm produced on each matter, then extended and pruned through attorney review at the firm'. The collections index lists Claude Opus 5 as the top model at 42.8%. The company says 10 of the 50 tasks have never been passed by any of the six models it evaluated at high reasoning effort: Claude Fable 5, Claude Opus 5, GPT-5.6 Sol, Gemini 3.6 Flash, Grok 4.5 and Kimi K3. Access is on request.
Why it matters: It extends Idler beyond retail into legal work, using data drawn from a real firm's work product.
- 13 Aug 2026benchmark
ShelfLife: a digital twin of a live e-commerce company, launched as the flagship benchmark
ShelfLife is a digital twin of a live, profitable multi-brand e-commerce company, built with the company's operators. The company presents it as 200 environments drawn from a decade of operational data, covering long-horizon planning, numerical finance precision, safety and complex tool use. Pass@1 on 2026-08-13: Claude Opus 5 69.0%, GPT-5.6 Sol 55.7%, Claude Fable 5 54.7%, Grok 4.6 39.4%, Kimi K3 30.6%, Gemini 3.6 Flash 16.1%. The YC launch claims that post-training Nemotron 3 Nano on 41 ShelfLife tasks raised its Finance Agent Benchmark score from 30.0% to 41.8% (p ≤ 0.001); this is a company claim.
Why it matters: It is the company's public launch artifact and offers both an eval and a claimed training signal (transfer to an external finance benchmark), which is what lab buyers look for.
Founding team
Ivan Chub
Co-founder & CEO
Previously co-founded Dark Forest, 0xPARC, Zupass and Hack Lodge. Worked at Facebook, Dynasty and AppSheet.
Tony Goss
Co-founder & CTO
Previously at Dark Forest DAO (@d_fdao) and 0xPARC, per his X bio. His YC bio says he '(ethically) hacked Dark Forest to win $400k'.
Nalu Concepcion
Co-founder
Leads the scaling of RL environments. Previously a software engineer at Microsoft, in product at Poshmark and in data science at PayMongo, and organized events of 500+ people.
What practitioners should know
For researchers
- ShelfLife, ShelfLife E-Sim and CorpLaw are published on idler.ai as leaderboards, with documentation and a sample task for ShelfLife. No paper, GitHub repo or Hugging Face dataset was found, and full datasets are available only on request.
- ShelfLife grading splits reward into quantitative (partial credit), integrity (penalties only, for specified harmful actions) and withholding (policy-violation deductions) criteria. The docs say models are not penalized for insufficient data, safety-motivated refusals, undisclosed rules or unexpected answer formatting.
- Data is deliberately left unsanitized. The docs say: 'We do not clean or alter schemas in the data workspace. We preserve the natural defects and gaps present in the raw data.'
- Company-reported training signal: Nemotron 3 Nano post-trained on 41 ShelfLife tasks rose from 30.0% to 41.8% on Finance Agent Benchmark (p ≤ 0.001), per the YC launch. No write-up or checkpoints were found.
For program managers
- YC S25, founded 2025, San Francisco. The YC profile listed a team size of 12 and 5 open roles (Forward Deployed Engineer, Design Engineer, Research Scientist, Software Engineer, Special Projects Operator) on 2026-10-05.
- The SaaS News (2026-08-21) reports a $9M seed led by Paradigm, with Y Combinator, Long Journey and angel investors. There is no primary announcement from the company or Paradigm, so treat it as reported.
- The company says 'the world's leading frontier labs' use its evals and environments but names none (self-claimed).
- Engagement is bespoke. Collections beyond the public leaderboards are 'Access Upon Request' via a scheduled call or an Airtable interest form. The Forward Deployed Engineer role 'own[s] the entire system that turns our partners' real work into evals & RL environments'.
- Environments are built from real partner businesses' data (a live retailer, a law firm) in collaboration with the people who did the work.
For data ops
- The ShelfLife sample shows MCP-style tools (read_file, list_files, read_policy, release_run, request_approval, whoami) and a /workspace/data directory of CSV exports (transfers, inventory snapshots, sales). Some tasks also give full terminal access.
- Some tasks run on a 'state-machine-driven business engine' that reproduces the sequence of information and events a human would meet, including irreversible decisions. An integrity violation (e.g. acting under an account that is not the agent's own) zeros the run.
- No public API, SDK, download or standard eval-harness format is documented. Access is arranged by request.
- No SOC 2, security page or trust center was found on idler.ai on 2026-10-05. Data is described as anonymized law-firm data or real retailer operational data.
Benchmarks & research
3 published
A long-horizon resource-management benchmark on a stateful simulator of retail operations, grounded in nearly a decade of real orders, purchase orders and inventory snapshots from a live retailer. Agents run the business one day at a time from a fixed date.
Grok 4.6 is top at 65.0 ± 1.7 mean reward points; Opus 5 scores 39.5.
A corporate-law benchmark built from a real law firm's anonymized data, set in transaction workspaces with documents, emails and client instructions. Grading criteria come from the firm's work product and were refined through attorney review. Access is on request.
Claude Opus 5 is top at 42.8% (per the collections index). The company says 10 of 50 tasks have never been passed by any of the six models evaluated at high reasoning effort.
A digital twin of a live, profitable multi-brand e-commerce company, built from a decade of operational data with its operators. The company presents it as 200 environments. Agents use MCP tools and/or full terminal access over raw CSV exports that are deliberately left uncleaned. Grading combines quantitative criteria (with partial credit), integrity criteria (deductions for harm) and withholding criteria (deductions for policy violations). A public leaderboard, documentation and a sample are published; the full dataset is on request.
Claude Opus 5 leads at 69.0% Pass@1, ahead of GPT-5.6 Sol (55.7%) and Claude Fable 5 (54.7%). Company-reported: training Nemotron 3 Nano on 41 ShelfLife tasks raised its Finance Agent Benchmark score from 30.0% to 41.8%.
Products
ShelfLife ↗A benchmark and environment set built as a digital twin of an e-commerce company. Public leaderboard, docs and sample; full data on request.launched 13 Aug 2026
ShelfLife E-Sim ↗A stateful, long-horizon retail-operations simulator.launched 26 Aug 2026
CorpLaw ↗A corporate legal practice benchmark built from anonymized law-firm data (access on request).launched 18 Aug 2026
Long-Horizon SWE collection ↗Repo tasks in Go, Python, Rust, C++, TS and Java. The company says each task has 100+ steps and a frontier-model pass rate below 50%. Access on request.
Terminal Agents collection ↗Execution-environment setup, enterprise API and tool use, and data cleaning and anonymization. Access on request.
Scale & velocity
Current headcount
12 (YC profile team size, as of 2026-10-05)
confirmed citeOpen roles
5 as of 2026-10-05 (Forward Deployed Engineer, Design Engineer, Research Scientist, Software Engineer, Special Projects Operator)
confirmed citeDistributed / remote
unknown
Research depth
Backgrounds
Ivan Chub (CEO): co-founder of Dark Forest, 0xPARC, Zupass and Hack Lodge; ex-Facebook, Dynasty, AppSheet, Tony Goss (CTO): previously Dark Forest DAO and 0xPARC; won $400k by (ethically) hacking Dark Forest, Nalu Concepcion (co-founder): ex-Microsoft SWE, Poshmark product, PayMongo data science
confirmed cite
Capital
Total raised
$9M+ (seed led by Paradigm, reported 2026-08; plus undisclosed YC S25 investment)
reported citeLast round
Seed, $9M, 2026-08-21, led by Paradigm with Y Combinator and Long Journey
reported citeInvestors
Paradigm (seed lead, reported), Y Combinator (S25)
reported cite| Date | Round | Amount | Led by | |
|---|
| 21 Aug 2026 | Seed | $9M | Paradigm | source ↗ |
| The SaaS News (2026-08-21) reports a $9M seed led by Paradigm, with Y Combinator, Long Journey and angel investors. There is no primary announcement from the company or Paradigm, so treat it as reported. |
| 2025 | Accelerator | undisclosed | Y Combinator | source ↗ |
| Accelerator (Y Combinator S25) · undisclosed |
Security & compliance
Other certifications
unknown
Product
What they sell
environments + evals
confirmed citeDeployment model
Bespoke/custom environments and private datasets delivered to labs on request. Public leaderboards, docs and samples are on idler.ai. The delivery mechanism is not publicly documented.
confirmed citeMaturity
early commercial (YC S25; public benchmarks launched Aug 2026)
estimated citeNotable customers
⚑Leading frontier AI labs (unnamed) self-claimed cite
Buyer analysis
Best fit: Frontier labs that want bespoke long-horizon RL environments and evals grounded in real company data (e-commerce operations, corporate law), or private long-horizon SWE and cybersecurity task sets.
How we verified this
2026-10-05 initial profile: I re-opened the YC profile, the idler.ai homepage, careers, the collections index, each benchmark page, the ShelfLife docs and sample, the blog, the Paradigm portfolio, the founders' and company's X bios. Identity is solid: three independent first-party signals tie idler.ai, YC S25 and the three founders together. All benchmark dates and leaderboard numbers matched, except that CorpLaw's 42.8% appears only on the collections index, so its source was changed. Several data-ops and summary details were tightened to match the docs' wording, and one unverifiable detail ('paginated plaintext') was removed. The founder headshots are YC-hosted JPEGs and were visually checked. The $9M Paradigm seed was later re-sourced to The SaaS News (2026-08-21) and stays "reported". Web-search quota was exhausted this session, so no new items (customers, SOC 2, other investors) could be discovered beyond direct fetches, and none were found on the company's own pages.
Sources
- idler.ai/ · 2026-10-05, Homepage: 'Reinforcement learning environments built from real production work'. No customers, investors or security mentions.
- idler.ai/careers · 2026-10-05, Self-description as a frontier data research lab used by leading frontier labs. 5 roles, all in San Francisco.
- www.ycombinator.com/companies/idler · 2026-10-05, YC S25 profile: founded 2025, team size 12, Active, SF, founder bios and headshots, 5 jobs, Nemotron 3 Nano / ShelfLife launch claim.
- idler.ai/collections · 2026-10-05, Collections index with task counts, measurement dates and top performers (including CorpLaw: Claude Opus 5 at 42.8%).
- idler.ai/collections/shelflife · 2026-10-05, ShelfLife leaderboard, measured 2026-08-13
- idler.ai/collections/shelflife/documentation · 2026-10-05, ShelfLife docs: MCP tools, terminal access, state-machine engine, reward decomposition, fairness safeguards, uncleaned data
- idler.ai/collections/shelflife/sample · 2026-10-05, Sample task: /workspace/data CSV exports; read_file/list_files/read_policy/release_run tools
- idler.ai/collections/shelflife-e-sim · 2026-10-05, ShelfLife E-Sim leaderboard, measured 2026-08-26
- idler.ai/collections/corplaw · 2026-10-05, CorpLaw description, measured 2026-08-18, access on request. Per-model scores are not shown on this page.
- idler.ai/collections/long-horizon-swe · 2026-10-05, Private Long-Horizon SWE collection
- idler.ai/collections/cybersecurity · 2026-10-05, Private cybersecurity collection
- idler.ai/collections/recursive-self-improvement · 2026-10-05, Private RSI collection
- idler.ai/collections/terminal-agents · 2026-10-05, Private terminal-agents collection
- thesaasnews.com/news/idler-raises-9m-seed · 2026-10-05, The SaaS News: idler raises 9M seed led by Paradigm (2026-08-21).
- www.paradigm.xyz/portfolio · 2026-10-05, Paradigm portfolio does not list Idler
- api.fxtwitter.com/chubivan · 2026-10-05, Ivan Chub X bio: cofounder of @idler_ai (yc s25), @0xPARC, @darkforest_eth, @hack_lodge
- api.fxtwitter.com/cha0sg0d_ · 2026-10-05, Tony Goss X bio: 'cto @idler_ai (YC S25) prev: @d_fdao | @0xParc'
- api.fxtwitter.com/idler_ai · 2026-10-05, Company X account: website idler.ai, joined 2025-06-04
- idler.ai/blog · 2026-10-05, Blog shows 'No posts yet'
- github.com/idler-ai · 2026-10-05, Name collision: unrelated personal data-analytics repos
Last updated 2026-10-05 · Every quantitative field carries a source and a confidence tag. Fields we could not source publicly are marked
unknown, never estimated. See the
methodology.