ProductsProduct / feature launchLarge Language Models · Efficiency & Inference1 source · Sep 30, 2026

axolotl-ai-cloud/axolotl v0.20.0

We have added a bunch of features including GGUF export, Ringmaster context parallelism, native NVFP4 LoRA, and expert parallelism without DeepEP.

Key points

  • You can now convert a trained checkpoint to GGUF for llama.cpp, Ollama, LM Studio, and llamafile with axolotl export config.yml --quantize Q4KM,Q80, which emits the base conversion plus one file per requested quant type.
  • You can now train long-context runs with packed sequences under Ulysses, Ring, or USP context parallelism, with experimental support for recurrent layers (GDN, KDA, Mamba2) that carry state across shards.
  • Setting contextparallelsize, or a contextparallel: block to pick the backend and Ulysses/Ring sizes, enables it without a plugins: entry; install with pip install axolotl[ringmaster]. examples/distributed-parallel/llama3-8b-ringmaster-cp.yaml pairs it with FSDP2.
  • You can now LoRA-train on a frozen native NVFP4 base under FSDP2, DeepSpeed ZeRO-1/2/3, or tensor parallelism, and the merged checkpoint matches what you trained: the forward runs on the same quantized base-plus-adapter weight that axolotl merge-lora writes.

Sources (1)

  • [1]axolotl-ai-cloud/axolotl v0.20.0
    GitHub: axolotl-ai-cloud/axolotl · Sep 30, 02:08 PM
    We have added a bunch of features including GGUF export, Ringmaster context parallelism, native NVFP4 LoRA, and expert parallelism without DeepEP.
    You can now convert a trained checkpoint to GGUF for llama.cpp, Ollama, LM Studio, and llamafile with `axolotl export config.yml --quantize Q4_K_M,Q8_0`, which emits the base conversion plus one file per requested quant type.

Extractive summary: sentences quoted from the sources.

Before this

  1. Sep 29, 2026NVIDIA/TensorRT-LLM v1.3.0rc29
  2. Aug 22, 2026sgl-project/sglang v0.5.18
  3. Jul 3, 2026huggingface/transformers v5.13.0: Release v5.13.0
  4. Jun 29, 2026vllm-project/vllm v0.24.0
  5. Jun 26, 2026sgl-project/sglang v0.5.14
  6. Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Related