791 Robot Datasets Converge
Cross-Embodiment Training Enters the Quality Governance Phase
This week scanned 86 HF orgs · 50 GitHub orgs · 71 blogs · 125 X accounts
LeRobot released 791 cross-embodiment community datasets on 2026-07-16, covering 46 robot categories and 25.97 million frames [P0]; Microsoft's Resource2Skill reached 2,098 downloads in its first week, expanding Agent training units from dialogue to executable skills [P0]; Meta's GAMUT evaluates 1,813 open-ended questions with a two-layer meta-rubric, and the best model scored only 58.7% [P1]. Strongest data demand signal this week: robotic manipulation trajectories.
Key Findings
This week's 5 high commercial value findings
lerobot/community_dataset_v3 was released on 2026-07-16. At the time of this scan, it had 5,114 downloads and 3 likes, with a scale of 10M<n<100M. The official data card disclosed that it was jointly built by 235 community contributors and includes 791 sub-datasets, 50,622 episodes, 25,971,082 frames, and 251.5 hours of data, covering 46 robot categories; of the original 851 datasets, 791 were retained after cleanup, with data removed due to missing videos, inconsistent data types, and failed conversions. Hugging Face also released Grabette on 2026-07-21, enabling participants to collect 6-DoF manipulation trajectories using a handheld gripper, camera, and browser workflow, then convert them into LeRobot format.
microsoft/RESOURCE2SKILL was released on 2026-07-17. At the time of this scan, it had 2,098 downloads and 9 likes. It covers workflows across Web, PowerPoint, Excel, Blender, and audio, distilling human-created tutorials, code, documentation, and reference outputs into discoverable, composable, executable skills. The project page reports that on the same GPT-5.4 backbone, average task scores across seven domains improved from 51.9 to 66.9, including a 38.2-point gain in UE5. During the same period, nvidia/Open-SWE-Traces was updated on 2026-07-16, containing 207,489 software engineering trajectories from SWE-agent/OpenHands, with 8,274 downloads at the time of this scan; Apple released a trajectory synthesis method on 2026-07-21 that does not require real API environments and filters results with an LLM judge.
facebook/GAMUT was released on 2026-07-20, with the accompanying paper published on 2026-07-21. This benchmark contains 1,813 questions based on real wearable images across 10 domains. Its method first uses a structured meta-rubric to describe the facts, alternatives, and ordered processes that a complete answer must satisfy, then mechanically compiles these into low-variance binary checklist items. The questions and rubrics were built through multiple rounds of a "human + LLM" process; each rubric is supported by web evidence and verified by experts. Among 14 closed- and open-source models, the best result was only 58.7%. AllenAI also released deepresearch-bench on 2026-07-22, providing 50 tasks each in Chinese and English and decomposing readability, insightfulness, completeness, and instruction following into weighted criteria.
surgeai/chartography completed this cycle's update on 2026-07-16 and had 121 downloads at the time of this scan. It contains only a test split, with 100 real charts and 100 questions spanning 12 professional domains including finance, healthcare, manufacturing, supply chain, chemistry, geoscience, and electrical engineering. Each question was written by relevant practitioners and reviewed by multiple experts, while also providing a complete calculation path and acceptable answer ranges defined by chart precision. The official report states that in the first round of evaluation, no frontier model exceeded 45%, and canary mechanisms were explicitly used to prevent training contamination. Amazon released ConfBench on 2026-07-20, generating up to 18 categories of degraded variants from 75 verified FCC invoices for a total of 1,346 samples, aimed at OCR robustness and confidence calibration.
nvidia/av-skills was updated on 2026-07-22. The official data card disclosed a total of 2,862,653 audio-video dialogues covering 29 task categories, with training videos up to 15 minutes long; among them, AV-Think contains 23,618 temporally anchored reasoning samples, with an average chain-of-thought length of 635.7 English words. The data covers speech, ambient sound, music, long-horizon video, and temporal localization. A ConsiSpace paper released on 2026-07-20 further points out that models tend to aggregate inconsistent spatial evidence under viewpoint changes and long-duration video, indicating that long-video capability has moved from clip-level recognition into a stage of cross-time, cross-view consistency reasoning.
Demand Signals
Infer training data demands from model releases
Want to discuss this issue?
Auto-generated by AI Dataset Radar · Updated weekly
AI Dataset Radar →