WorldBench: Evaluating LLMs on Three.js Voxel World Generation
We present WorldBench, a benchmark and judge for open-ended, LLM-generated Three.js worlds.
ProofPaper ↗
Key points
- Large language models can now write complete, interactive 3D worlds as code, but grading those worlds automatically is unreliable.
- On worlds written by five frontier models we find that the two views disagree on 32% of required items, mostly code that no frame shows, and that fixed views miss small close-up contents.
- From one prompt describing a floating voxel island with ten biomes, physics, and day/night and seasonal cycles, the judge explores the running world, controlling its clock, orbiting it, and sending a navigator agent to frame each biome, and reads the code for what it sees.
- We evaluate five frontier models: Claude Fable 5.1, GPT-6 Astra, Kimi K3, Grok 4.7 and Gemini 3.1 Pro.
Sources (1)
- [1]WorldBench: Evaluating LLMs on Three.js Voxel World GenerationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:52 AM
We present WorldBench, a benchmark and judge for open-ended, LLM-generated Three.js worlds.
Large language models can now write complete, interactive 3D worlds as code, but grading those worlds automatically is unreliable.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
- Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
- Oct 3, 2026[AINews] not much happened today
- Oct 1, 2026OpenAI Security: Controlling Models is Now ‘Hell’
- Oct 1, 2026The Dot and the Swarm
- Sep 30, 2026pydantic/pydantic-ai v2.52.0: v2.52.0 (2026-09-29)