FetchMan-data in October
AI Data Intelligence Weekly
This week scanned 86 HF orgs · 50 GitHub orgs · 71 blogs · 125 X accounts
FetchMan-data released 52 million embodied interaction frames on October 1 [P0]; Surge AI launched the GDP.xlsx benchmark for professional judgment on September 30 [P0]; and an October 1 paper elevated the “provenance and audit trail” of synthetic data to a training data infrastructure issue [P1]. The strongest data demand signal this week: continuous embodied intelligence trajectories.
Key Findings
This week's 5 high commercial value findings
Allen AI released allenai/fetchman-data on 2026-10-01, containing 352,696 robot training episodes, 52,008,151 frames, and 90 language tasks. The dataset covers actions, first-person video, and proprioceptive information from a Unitree G1 humanoid robot performing object-grasping tasks in MolmoSpaces; as of the scan date, it had 434 downloads and 4 likes.
Surge AI released surgeai/GDP.xlsx on 2026-09-30, with 71 downloads as of the scan date. The benchmark contains 70 tasks across 13 professional domains and is packaged in the Harbor environment to evaluate whether AI Agents can understand spreadsheets and provide answers consistent with the judgment of domain professionals.
The paper Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects was published on 2026-10-01. It proposes binding synthetic speech research objects to source specifications, generated content, waveforms, targets, factual requirements, quality signals, and audit lineage. The paper LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification, also published that day, explores using verifiable rewards to improve time-series captioning quality.
InternLM released internlm/AdvancedMathBench on 2026-09-29, with 72 downloads and 1 like. It covers natural-language mathematical proof generation and verification at the undergraduate and qualification-exam levels. The companion model internlm/AdvancedMathBench-AutoVerifier was released the same day, with 50 downloads and 1 like. Intern-Decision-4B, Intern-Decision-2B, and Intern-Decision-0.8B also recorded 1,607, 575, and 1,159 downloads, respectively, and are designed for multimodal decision-making and structured prediction.
Meta released facebook/adepts on 2026-09-28, with 133 downloads and 1 like as of the scan date. ADEPTS-BENCH uses dual-stream reliability evaluation to assess the ability of mobile and desktop Computer Use Agents to handle ambiguity, execute operations, and ensure safety in visual interfaces.
Demand Signals
Infer training data demands from model releases
Want to discuss this issue?
Auto-generated by AI Dataset Radar · Updated weekly
AI Dataset Radar →