Radar Brief Week 24, 2026 · 2026-07-16 — 2026-07-23

791 Robot Datasets Converge
Cross-Embodiment Training Enters the Quality Governance Phase

This week scanned 86 HF orgs · 50 GitHub orgs · 71 blogs · 125 X accounts

0
Valuable Datasets
0
Related Papers
0
Blog Posts
0
Active Repos
One-line Summary

LeRobot released 791 cross-embodiment community datasets on 2026-07-16, covering 46 robot categories and 25.97 million frames [P0]; Microsoft's Resource2Skill reached 2,098 downloads in its first week, expanding Agent training units from dialogue to executable skills [P0]; Meta's GAMUT evaluates 1,813 open-ended questions with a two-layer meta-rubric, and the best model scored only 58.7% [P1]. Strongest data demand signal this week: robotic manipulation trajectories.

Key Findings

This week's 5 high commercial value findings

P0 LeRobot released 791 cross-embodiment community datasets on 2026-07-16, covering 46 robot categories and 25.97 million frames [P0]

lerobot/community_dataset_v3 was released on 2026-07-16. At the time of this scan, it had 5,114 downloads and 3 likes, with a scale of 10M<n<100M. The official data card disclosed that it was jointly built by 235 community contributors and includes 791 sub-datasets, 50,622 episodes, 25,971,082 frames, and 251.5 hours of data, covering 46 robot categories; of the original 851 datasets, 791 were retained after cleanup, with data removed due to missing videos, inconsistent data types, and failed conversions. Hugging Face also released Grabette on 2026-07-21, enabling participants to collect 6-DoF manipulation trajectories using a handheld gripper, camera, and browser workflow, then convert them into LeRobot format.

Business impact → Robot data collection is moving from single-lab, single-platform setups toward community-driven, cross-embodiment convergence. The real bottleneck is shifting from "whether data exists" to "whether heterogeneous devices, viewpoints, frame rates, and task semantics can be merged reliably." This requires human judgment on task descriptions, trajectory completeness, operation success or failure, anomaly recovery, and safety boundaries. Knowlyr can prioritize building capabilities for cross-embodiment data quality inspection, episode-level success determination for robotics, layered failure taxonomy, and disputed trajectory review—turning community contributions into trainable data assets.
P0 Microsoft's Resource2Skill reached 2,098 downloads in its first week, expanding Agent training units from dialogue to executable skills [P0]

microsoft/RESOURCE2SKILL was released on 2026-07-17. At the time of this scan, it had 2,098 downloads and 9 likes. It covers workflows across Web, PowerPoint, Excel, Blender, and audio, distilling human-created tutorials, code, documentation, and reference outputs into discoverable, composable, executable skills. The project page reports that on the same GPT-5.4 backbone, average task scores across seven domains improved from 51.9 to 66.9, including a 38.2-point gain in UE5. During the same period, nvidia/Open-SWE-Traces was updated on 2026-07-16, containing 207,489 software engineering trajectories from SWE-agent/OpenHands, with 8,274 downloads at the time of this scan; Apple released a trajectory synthesis method on 2026-07-21 that does not require real API environments and filters results with an LLM judge.

Business impact → The basic unit of Agent data is shifting from "instruction–response" to execution chains of "requirement–skill–tool call–environment feedback–verifiable result." Synthetic trajectories will scale rapidly, but whether they reflect real needs, whether steps are necessary, whether tool calls are reliable, and whether outputs meet professional standards still depend on high-quality judgment. Knowlyr should further productize Agent trajectory review into five judgment tasks: skill extraction, task validity verification, process preference comparison, failure attribution, and real software acceptance.
P1 Meta's GAMUT evaluates 1,813 open-ended questions with a two-layer meta-rubric, and the best model scored only 58.7% [P1]

facebook/GAMUT was released on 2026-07-20, with the accompanying paper published on 2026-07-21. This benchmark contains 1,813 questions based on real wearable images across 10 domains. Its method first uses a structured meta-rubric to describe the facts, alternatives, and ordered processes that a complete answer must satisfy, then mechanically compiles these into low-variance binary checklist items. The questions and rubrics were built through multiple rounds of a "human + LLM" process; each rubric is supported by web evidence and verified by experts. Among 14 closed- and open-source models, the best result was only 58.7%. AllenAI also released deepresearch-bench on 2026-07-22, providing 50 tasks each in Chinese and English and decomposing readability, insightfulness, completeness, and instruction following into weighted criteria.

Business impact → Evaluation of open-ended generation is evolving from "is the answer correct" to "is it complete, is the evidence sufficient, and does the process satisfy domain structure." Rubric design itself is becoming a scarce data asset rather than an auxiliary note. Knowlyr can provide evidence-driven rubric writing, criterion-level scoring, expert arbitration, and reviewer consistency calibration—preserving human judgment as reusable, auditable evaluation infrastructure.
P1 Surge AI's Chartography uses professionals to build 100 real chart questions, with frontier models scoring below 45% at best [P1]

surgeai/chartography completed this cycle's update on 2026-07-16 and had 121 downloads at the time of this scan. It contains only a test split, with 100 real charts and 100 questions spanning 12 professional domains including finance, healthcare, manufacturing, supply chain, chemistry, geoscience, and electrical engineering. Each question was written by relevant practitioners and reviewed by multiple experts, while also providing a complete calculation path and acceptable answer ranges defined by chart precision. The official report states that in the first round of evaluation, no frontier model exceeded 45%, and canary mechanisms were explicitly used to prevent training contamination. Amazon released ConfBench on 2026-07-20, generating up to 18 categories of degraded variants from 75 verified FCC invoices for a total of 1,346 samples, aimed at OCR robustness and confidence calibration.

Business impact → The challenge in professional visual reasoning is not label recognition, but understanding sparse coordinates, overlapping flows, projection geometry, and industry conventions not stated in the prompt. An answer that appears plausible but misreads a curve may fail outright in healthcare, engineering, or finance. Knowlyr can organize professionals to build golden evaluation sets with "real materials + calculation paths + tolerance ranges + anti-contamination mechanisms," converting domain expertise into measurable judgment standards.
P2 NVIDIA expands AV-Skills to 2.86 million audio-video dialogues as long-video training shifts toward temporal localization and cross-modal consistency [P2]

nvidia/av-skills was updated on 2026-07-22. The official data card disclosed a total of 2,862,653 audio-video dialogues covering 29 task categories, with training videos up to 15 minutes long; among them, AV-Think contains 23,618 temporally anchored reasoning samples, with an average chain-of-thought length of 635.7 English words. The data covers speech, ambient sound, music, long-horizon video, and temporal localization. A ConsiSpace paper released on 2026-07-20 further points out that models tend to aggregate inconsistent spatial evidence under viewpoint changes and long-duration video, indicating that long-video capability has moved from clip-level recognition into a stage of cross-time, cross-view consistency reasoning.

Business impact → A scale of 2.86 million can be expanded rapidly by machines, but whether audio and video are aligned, whether event boundaries are accurate, whether reasoning cites the correct time span, and whether cross-view spatial relations are consistent still require traceable judgment standards. Knowlyr can invest in long-video event segmentation, temporal evidence localization, audio-video consistency review, spatial relation adjudication, and reasoning-to-evidence alignment—forming high-value multimodal judgment capabilities distinct from low-barrier content production.

Demand Signals

Infer training data demands from model releases

Data Type Intensity Trend Related Signals
Robotic Manipulation Trajectories
Very High → Continuing
LeRobot aggregates 791 datasets · 50,622 episodes · 46 robot categories; Grabette lowers the barrier for 6-DoF demonstration collection
Agent Real-World Trajectories and Long-Horizon Evaluation
Very High → Continuing
Open-SWE-Traces released 207,489 trajectories and reached 8,274 downloads; Resource2Skill reached 2,098 downloads
Meta-Rubrics for Open-Ended Generation
Very High ↑ New
GAMUT validates two-layer meta-rubrics with 1,813 questions, with the best model scoring only 58.7%; AllenAI simultaneously released Chinese and English criteria sets
Evaluation Data for Scientific Analysis
High → Continuing
Chartography covers 12 professional domains and frontier models score below 45%; SciForma emphasizes structural fidelity in scientific diagrams
Egocentric Video Memory / 3D Reasoning
High → Continuing
Open-AoE releases about 2,000 hours of egocentric operation video; Grabette and ConsiSpace concurrently reinforce trajectory and spatial consistency
Temporal Localization in Long Audio-Video
High ↑ New
AV-Skills expands to 2,862,653 dialogues, including 23,618 temporally anchored reasoning samples
Executable Skill Libraries for Agents
High ↑ New
Resource2Skill covers seven categories of real software tasks, with the project reporting a 15.0-point average task score improvement
Synthetic Environments and World Model Data
High ↑ New
Apple generates stateful trajectories from API specifications; ABot-World-0 and G-MAD expand interactive worlds and RGB-T simulation data respectively
Document Robustness and Confidence Calibration
Medium ↑ New
ConfBench generates 1,346 degraded samples from 75 verified invoices, covering up to 18 categories of real noise
Expert Golden Sets and Anti-Contamination Evaluation
Medium ↑ New
Chartography uses professional authorship · multi-round review · answer tolerances and canary anti-contamination
Speech Multilingual QA Evaluation ↓ Dropped Present in previous issue, absent this issue
Retrieval Citation and Scientific Agent Evidence Data ↓ Dropped Present in previous issue, absent this issue
Preference Learning and Noise Correction Data ↓ Dropped Present in previous issue, absent this issue
Data Attribution and Model Memory Analysis ↓ Dropped Present in previous issue, absent this issue
Privacy and Synthetic Sensitive Text ↓ Dropped Present in previous issue, absent this issue
Climate / Simulation Temporal Environments ↓ Dropped Present in previous issue, absent this issue

Want to discuss this issue?

Kai
Kai Founder & CEO
苏文
苏文 AI Documentation & Release Engineer
陆明哲
陆明哲 AI Product Manager

Auto-generated by AI Dataset Radar · Updated weekly

AI Dataset Radar →