ResearchResearch paperImage, Video & 3D Generation · Large Language Models · Robotics & Embodied AI1 source · Oct 7, 2026

WorldBench: Evaluating LLMs on Three.js Voxel World Generation

We present WorldBench, a benchmark and judge for open-ended, LLM-generated Three.js worlds.

Key points

  • Large language models can now write complete, interactive 3D worlds as code, but grading those worlds automatically is unreliable.
  • On worlds written by five frontier models we find that the two views disagree on 32% of required items, mostly code that no frame shows, and that fixed views miss small close-up contents.
  • From one prompt describing a floating voxel island with ten biomes, physics, and day/night and seasonal cycles, the judge explores the running world, controlling its clock, orbiting it, and sending a navigator agent to frame each biome, and reads the code for what it sees.
  • We evaluate five frontier models: Claude Fable 5.1, GPT-6 Astra, Kimi K3, Grok 4.7 and Gemini 3.1 Pro.

Sources (1)

  • [1]WorldBench: Evaluating LLMs on Three.js Voxel World Generation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:52 AM
    We present WorldBench, a benchmark and judge for open-ended, LLM-generated Three.js worlds.
    Large language models can now write complete, interactive 3D worlds as code, but grading those worlds automatically is unreliable.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
  2. Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
  3. Oct 3, 2026[AINews] not much happened today
  4. Oct 1, 2026OpenAI Security: Controlling Models is Now ‘Hell’
  5. Oct 1, 2026The Dot and the Swarm
  6. Sep 30, 2026pydantic/pydantic-ai v2.52.0: v2.52.0 (2026-09-29)

Related