microsoft/AesCode 8B and 32B
AesCode generates information-rich visual artifacts such as slides, posters, and dashboards as HTML/CSS.
Key points
- The output remains structured, editable, and verifiable, but the task poses a distinct challenge: code models cannot see how layout, hierarchy, and color come together on the canvas.
- AesCode uses an image generated from the same prompt as an aesthetic reference while following the prompt for the required content.
- AesCode separates semantic requirements from visual cues through graph-structured supervision and decoupled cross-modal rewards.
- AesCode-32B starts from Qwen3-VL-32B-Instruct and is trained with cold-start SFT followed by GDPO across seven reward channels.
Sources (1)
- [1]microsoft/AesCode 8B and 32Br/LocalLLaMA (top, daily) · Oct 11, 11:17 AM
AesCode generates information-rich visual artifacts such as slides, posters, and dashboards as HTML/CSS.
The output remains structured, editable, and verifiable, but the task poses a distinct challenge: code models cannot see how layout, hierarchy, and color come together on the canvas.
Extractive summary: sentences quoted from the sources.