DataSense-Bench: The First Step Toward an AI Scientist
We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning.
Key points
- As claims about recursive self-improvement (RSI) and artificial general intelligence (AGI) proliferate, we ask a simple question: do frontier AI models have a sense of data, i.e., can they reliably select the right data for training?
- We ask AI agents to select and rank candidate training subsets that can be used to fine-tune a small LLM model.
- We instantiate the benchmark in terminal problem solving and tool use, selecting trajectories from OpenThoughts-Agent and EnvScaler and evaluating on TBLite and BFCL, respectively.
- In our experiments, selection gains over random selection are limited; agents do not reliably rank their selected groups, and ranking ability does not hold consistently across tasks: Astra identifies the best group in all three tool-use runs but in only one of three terminal runs.
Sources (1)
- [1]DataSense-Bench: The First Step Toward an AI ScientistarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:50 PM
We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning.
As claims about recursive self-improvement (RSI) and artificial general intelligence (AGI) proliferate, we ask a simple question: do frontier AI models have a sense of data, i.e., can they reliably select the right data for training?
Extractive summary: sentences quoted from the sources.