AION
Modelauto-detected

meta-llama/Llama-3.1-8B

Also known as: Llama-3.1-8B

2stories this week
2last 30 days
2all time

In the model registry

ModelParamsContextReleased
meta-llama/Llama-3.1-8B8B–Jul 14, 2024

Timeline

  1. Oct 8, 2026 · Research paper · 2 sources
    SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
    The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency.
  2. Oct 6, 2026 · Research paper · 2 sources
    CARE: Certifying Acceleration for Vision-Language-Action Inference
    Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success.

Often appears with