axolotl-ai-cloud/axolotl v0.20.0
We have added a bunch of features including GGUF export, Ringmaster context parallelism, native NVFP4 LoRA, and expert parallelism without DeepEP.
ProofCode ↗
Key points
- You can now convert a trained checkpoint to GGUF for llama.cpp, Ollama, LM Studio, and llamafile with axolotl export config.yml --quantize Q4KM,Q80, which emits the base conversion plus one file per requested quant type.
- You can now train long-context runs with packed sequences under Ulysses, Ring, or USP context parallelism, with experimental support for recurrent layers (GDN, KDA, Mamba2) that carry state across shards.
- Setting contextparallelsize, or a contextparallel: block to pick the backend and Ulysses/Ring sizes, enables it without a plugins: entry; install with pip install axolotl[ringmaster]. examples/distributed-parallel/llama3-8b-ringmaster-cp.yaml pairs it with FSDP2.
- You can now LoRA-train on a frozen native NVFP4 base under FSDP2, DeepSpeed ZeRO-1/2/3, or tensor parallelism, and the merged checkpoint matches what you trained: the forward runs on the same quantized base-plus-adapter weight that axolotl merge-lora writes.
Sources (1)
- [1]axolotl-ai-cloud/axolotl v0.20.0GitHub: axolotl-ai-cloud/axolotl · Sep 30, 02:08 PM
We have added a bunch of features including GGUF export, Ringmaster context parallelism, native NVFP4 LoRA, and expert parallelism without DeepEP.
You can now convert a trained checkpoint to GGUF for llama.cpp, Ollama, LM Studio, and llamafile with `axolotl export config.yml --quantize Q4_K_M,Q8_0`, which emits the base conversion plus one file per requested quant type.
Extractive summary: sentences quoted from the sources.
Before this
- Sep 29, 2026NVIDIA/TensorRT-LLM v1.3.0rc29
- Aug 22, 2026sgl-project/sglang v0.5.18
- Jul 3, 2026huggingface/transformers v5.13.0: Release v5.13.0
- Jun 29, 2026vllm-project/vllm v0.24.0
- Jun 26, 2026sgl-project/sglang v0.5.14
- Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model

