Radar Brief Week 31, 2026 · 2026-09-26 — 2026-10-03

FetchMan-data in October
AI Data Intelligence Weekly

This week scanned 86 HF orgs · 50 GitHub orgs · 71 blogs · 125 X accounts

0
Valuable Datasets
0
Related Papers
0
Blog Posts
0
Active Repos
One-line Summary

FetchMan-data released 52 million embodied interaction frames on October 1 [P0]; Surge AI launched the GDP.xlsx benchmark for professional judgment on September 30 [P0]; and an October 1 paper elevated the “provenance and audit trail” of synthetic data to a training data infrastructure issue [P1]. The strongest data demand signal this week: continuous embodied intelligence trajectories.

Key Findings

This week's 5 high commercial value findings

P0 FetchMan-data released 52 million embodied interaction frames on October 1 [P0]

Allen AI released allenai/fetchman-data on 2026-10-01, containing 352,696 robot training episodes, 52,008,151 frames, and 90 language tasks. The dataset covers actions, first-person video, and proprioceptive information from a Unitree G1 humanoid robot performing object-grasping tasks in MolmoSpaces; as of the scan date, it had 434 downloads and 4 likes.

Business significance → Embodied intelligence training is moving from static images and single-step actions toward long-horizon data consisting of “language instruction—continuous action—visual feedback—body state.” Scalable synthetic data solves the quantity problem, but successful grasping in the real world, occlusion handling, object fragility, and action safety still require human judgment to define valid behavior and failure boundaries. Knowlyr should prioritize robot trajectory review, anomalous action filtering, task success judgment, and sim-to-real transfer evaluation, enabling contributors who provide judgment to participate directly in embodied model development and earn income.
P0 Surge AI launched the GDP.xlsx benchmark for professional judgment on September 30 [P0]

Surge AI released surgeai/GDP.xlsx on 2026-09-30, with 71 downloads as of the scan date. The benchmark contains 70 tasks across 13 professional domains and is packaged in the Harbor environment to evaluate whether AI Agents can understand spreadsheets and provide answers consistent with the judgment of domain professionals.

Business significance → The value of training data is shifting from “can it generate correct text?” to “can it make professional decisions using real-world work materials?” Spreadsheet understanding, cross-cell reasoning, anomaly identification, and business conclusion judgment are difficult to validate solely through automated rules. Professionals’ judgment processes, confidence levels, and acceptable error margins will become high-value data assets. Knowlyr can break down complex spreadsheet tasks in finance, operations, supply chain, legal, and healthcare into paid judgment contributions, creating data products for professional Agent training and evaluation.
P1 An October 1 paper elevated the “provenance and audit trail” of synthetic data to a training data infrastructure issue [P1]

The paper Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects was published on 2026-10-01. It proposes binding synthetic speech research objects to source specifications, generated content, waveforms, targets, factual requirements, quality signals, and audit lineage. The paper LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification, also published that day, explores using verifiable rewards to improve time-series captioning quality.

Business significance → Synthetic data is shifting from one-off content generation toward traceable, auditable, and verifiable training objects. Customers need not only the results, but also to know which models generated the data, whose judgment was applied, which quality signals were confirmed, and which samples should be discarded. Knowlyr can build a judgment contribution chain covering “generation records—human review—multi-round dispute resolution—version tracking,” with a focus on speech, time series, and multimodal training data to reduce the risk of unexplainability after synthetic data enters production training.
P1 InternLM released AdvancedMathBench and AutoVerifier on September 29, strengthening proof generation and verification [P1]

InternLM released internlm/AdvancedMathBench on 2026-09-29, with 72 downloads and 1 like. It covers natural-language mathematical proof generation and verification at the undergraduate and qualification-exam levels. The companion model internlm/AdvancedMathBench-AutoVerifier was released the same day, with 50 downloads and 1 like. Intern-Decision-4B, Intern-Decision-2B, and Intern-Decision-0.8B also recorded 1,607, 575, and 1,159 downloads, respectively, and are designed for multimodal decision-making and structured prediction.

Business significance → Model training is shifting from “generating answers” to “generating decisions and proving why those decisions were made.” Automated verifiers can check formal correctness, but cannot fully replace human judgment on whether a proof is meaningful, whether the reasoning conforms to educational or business standards, or whether the conclusion is acceptable. Knowlyr can focus on mathematical proof quality, completeness of decision rationales, multimodal evidence consistency, and refusal boundaries to establish high-difficulty reasoning data services combining “machine verification + human review.”
P2 Meta released ADEPTS on September 28 to evaluate the reliability of cross-device Computer Use Agents [P2]

Meta released facebook/adepts on 2026-09-28, with 133 downloads and 1 like as of the scan date. ADEPTS-BENCH uses dual-stream reliability evaluation to assess the ability of mobile and desktop Computer Use Agents to handle ambiguity, execute operations, and ensure safety in visual interfaces.

Business significance → The core risks of Computer Use Agents are no longer limited to incorrect answers; they also include clicking the wrong button, misunderstanding interface states, bypassing security prompts, or continuing execution in ambiguous scenarios. These tasks require human judgment about “whether to continue, whether clarification is needed, and which step constitutes a high-risk action,” driving demand for GUI operation trajectories, risk levels, stopping conditions, and corrective feedback data. Knowlyr can package critical operation judgments across websites, desktop software, and mobile applications as high-value safety evaluation tasks.

Demand Signals

Infer training data demands from model releases

Data Type Intensity Trend Related Signals
Continuous embodied intelligence trajectories
Extremely strong ↑ New
Allen AI **fetchman-data** released 352,696 episodes · 52 million frames, covering language instructions · actions · video and proprioception
Multimodal spatial and 3D interaction
Extremely strong ↑ New
Meta **show3d-dataset** reached 23,202 downloads; Google **polaris-bench** focuses on polar-coordinate visual reasoning; NVIDIA continues to release Physical AI data
Agent tool calling and long-horizon trajectories
Extremely strong ↑ New
NVIDIA **Nemotron-RL-Agentic-Function-Calling-Pivot-v1** · SWE Pivot · Conversational Tool Use Pivot reached 1,802 · 1,437 and 1,310 downloads, respectively
Computer Use safety and GUI operations
Strong ↑ New
Meta **ADEPTS** was released on 2026-09-28, covering mobile · desktop · ambiguity handling and reliability
Professional spreadsheets and enterprise knowledge judgment
Strong ↑ New
Surge **GDP.xlsx** contains 70 tasks · 13 professional domains; discussions of enterprise dirty knowledge-base retrieval benchmarks also appeared on HN during the same period
RLHF and preference calibration
Extremely strong ↑ New
NVIDIA **Nemotron-RL-Safety-v1** · Allen AI AstaBrief DPO/SFT data, alongside **Optimal Design for Active Preference Learning with Biased LLM Judges**, advanced during the same period
Reward signals in legal, medical, and other domains
Strong ↑ New
**LexReward** focuses on multidimensional rewards for legal language models; Snorkel **MedPAIR** showed that agreement between doctors and models on relevance judgment was only 50%—60%
Mathematical proof and decision verification
Strong ↑ New
InternLM **AdvancedMathBench** and **AutoVerifier** were released on 2026-09-29, covering proof generation and verification
Synthetic data provenance and quality auditing
Strong ↑ New
**Generation Provenance Before Behavior Attribution** · **Effective Synthetic Data Curation Requires Group-Level Signals** emphasize generation provenance · audit lineage and group-level filtering
Video anomaly and safety incident understanding
Strong ↑ New
NVIDIA **PhysicalAI-Event-Videos** contains 3,796 labeled segments and 71,023 description queries; **PhysicalAI-Traffic-Anomaly-Reasoning** targets traffic anomaly reasoning
Code Agents and software engineering trajectories
Extremely strong ↑ New
NVIDIA SWE Pivot dataset · OpenAI Codex has 127,645 stars on GitHub · Anthropic Claude Code has 148,983 stars, continuing to validate demand for Code Agent training
Multilingual speech and dialect judgment
Moderate ↑ New
SpeechOcean Dolphin supports 40 Eastern languages and 22 Chinese dialects; interest in the Eleven v4/Turbo products is also rising
Software engineering Agent trajectories ↓ Dropped Present in previous issue, absent this issue
Enterprise knowledge work judgment ↓ Dropped Present in previous issue, absent this issue
Agent decision-making and verification data ↓ Dropped Present in previous issue, absent this issue
Computer use and cross-system operations ↓ Dropped Present in previous issue, absent this issue
Real-world robot interaction data ↓ Dropped Present in previous issue, absent this issue
Multimodal 3D and spatial data ↓ Dropped Present in previous issue, absent this issue
Speech intent, multilingual dialects, and escalation risk ↓ Dropped Present in previous issue, absent this issue
Reward models and Reward Hacking data ↓ Dropped Present in previous issue, absent this issue
Preference learning and continuous feedback ↓ Dropped Present in previous issue, absent this issue
Synthetic edge cases and data acceptance ↓ Dropped Present in previous issue, absent this issue
Scientific RAG evidence chains ↓ Dropped Present in previous issue, absent this issue

Want to discuss this issue?

Kai
Kai Founder & CEO
苏文
苏文 AI Documentation & Release Engineer
陆明哲
陆明哲 AI Product Manager

Auto-generated by AI Dataset Radar · Updated weekly

AI Dataset Radar →