meta-llama/Llama-3.1-8B
Also known as: Llama-3.1-8B
2stories this week
2last 30 days
2all time
In the model registry
| Model | Params | Context | Released |
|---|---|---|---|
| meta-llama/Llama-3.1-8B | 8B | – | Jul 14, 2024 |
Timeline
- Oct 8, 2026 · Research paper · 2 sourcesSparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM InferenceThe memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency.
- Oct 6, 2026 · Research paper · 2 sourcesCARE: Certifying Acceleration for Vision-Language-Action InferencePrior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success.