4-Tensor Attention Model for Semantic Physical Reality
We describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning.
Key points
- A window of states has positions (x, t) and two fibers, a semantic fiber and a temporal-context fiber, and one softmax normalizes attention jointly over the window.
- To test the update on its own, we train on ROCStories, where each window poses the same next-sentence task at the semantic layer.
- On the validation split, with one seed per setting, the last-sentence cross-entropy on the three matched settings is lower for the 4-tensor model than for a free-running one-dimensional transformer by 5.3% at H=2, L=2, by 2.6% at H=4, L=2, and by 2.4% at H=4, L=3.
- On the same two GPUs that 4-tensor run finished in 2.4 hours and the baseline run in 45.2 hours; the baseline is trained by free-running decoding, one sequential forward pass per target token.
Sources (1)
- [1]4-Tensor Attention Model for Semantic Physical RealityarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 11:19 AM
We describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning.
A window of states has positions (x, t) and two fibers, a semantic fiber and a temporal-context fiber, and one softmax normalizes attention jointly over the window.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
- Oct 8, 2026One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
- Oct 8, 2026Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0