AION
Research paperTraining & Scaling · Robotics & Embodied AI · Large Language Models1 source · Oct 8, 2026

DataSense-Bench: The First Step Toward an AI Scientist

We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning.

Key points

  • As claims about recursive self-improvement (RSI) and artificial general intelligence (AGI) proliferate, we ask a simple question: do frontier AI models have a sense of data, i.e., can they reliably select the right data for training?
  • We ask AI agents to select and rank candidate training subsets that can be used to fine-tune a small LLM model.
  • We instantiate the benchmark in terminal problem solving and tool use, selecting trajectories from OpenThoughts-Agent and EnvScaler and evaluating on TBLite and BFCL, respectively.
  • In our experiments, selection gains over random selection are limited; agents do not reliably rank their selected groups, and ranking ability does not hold consistently across tasks: Astra identifies the best group in all three tool-use runs but in only one of three terminal runs.

Sources (1)

  • [1]DataSense-Bench: The First Step Toward an AI Scientist
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:50 PM
    We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning.
    As claims about recursive self-improvement (RSI) and artificial general intelligence (AGI) proliferate, we ask a simple question: do frontier AI models have a sense of data, i.e., can they reliably select the right data for training?

Extractive summary: sentences quoted from the sources.