Open weights
Locally runnable models
Explore downloadable models by architecture, parameter count, published context, licence and recorded weight formats. Specifications do not establish whether a model will fit or run well on a particular computer.
Catalogue review recorded · individual claims have separate datesStructured datasetSourceMethodology
A context ceiling is not a retrieval-quality result. For mixture-of-experts models, active parameters describe the subset used during a token step; they are not the total weights that must be stored. No model-fit verdict is inferred from these specifications.
Llama 3.1 8B Instruct
8.0B total · Llama 3.1 Community License · 131,072 context tokens
Architecture, licence and source
GQA 8 KV heads. Weight memory uses 8.03B total parameters.
Attention: GQA · layers: 32 · KV heads: 8. Architecture evidence: architecture.
Licence: Llama 3.1 Community License. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2024-07-23 · source review recorded 2026-09-04.
Publisher model card and weightsLlama 3.1 70B Instruct
70.6B total · Llama 3.1 Community License · 131,072 context tokens
Architecture, licence and source
No additional editorial note recorded.
Attention: GQA · layers: 80 · KV heads: 8. Architecture evidence: architecture.
Licence: Llama 3.1 Community License. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2024-07-23 · source review recorded 2026-09-04.
Publisher model card and weightsLlama 3.3 70B Instruct
70.6B total · Llama 3.3 Community License · 131,072 context tokens
Architecture, licence and source
Same 70B GQA layout as Llama 3.1 70B; instruction weights differ.
Attention: GQA · layers: 80 · KV heads: 8. Architecture evidence: architecture.
Licence: Llama 3.3 Community License. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2024-12-06 · source review recorded 2026-09-04.
Publisher model card and weightsQwen2.5 7B Instruct
7.6B total · Apache 2.0 · 131,072 context tokens
Architecture, licence and source
Native 32K; 128K via YaRN as documented on the model card.
Attention: GQA · layers: 28 · KV heads: 4. Architecture evidence: architecture.
Licence: Apache 2.0. Commercial-use claim: verified. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2024-09-19 · source review recorded 2026-09-04.
Publisher model card and weightsQwen2.5 32B Instruct
32.5B total · Apache 2.0 · 131,072 context tokens
Architecture, licence and source
No additional editorial note recorded.
Attention: GQA · layers: 64 · KV heads: 8. Architecture evidence: architecture.
Licence: Apache 2.0. Commercial-use claim: verified. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2024-09-19 · source review recorded 2026-09-04.
Publisher model card and weightsQwen2.5 72B Instruct
72.7B total · Qwen License · 131,072 context tokens
Architecture, licence and source
Tongyi Qianwen license — not Apache 2.0. Commercial use follows the published license.
Attention: GQA · layers: 80 · KV heads: 8. Architecture evidence: architecture.
Licence: Qwen License. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2024-09-19 · source review recorded 2026-09-04.
Publisher model card and weightsQwen2.5-Coder 32B Instruct
32.5B total · Apache 2.0 · 131,072 context tokens
Architecture, licence and source
No additional editorial note recorded.
Attention: GQA · layers: 64 · KV heads: 8. Architecture evidence: architecture.
Licence: Apache 2.0. Commercial-use claim: verified. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2024-11-11 · source review recorded 2026-09-04.
Publisher model card and weightsDeepSeek-V3
671B total · DeepSeek Model License · 131,072 context tokens
Architecture, licence and source
MoE 671B total / 37B active. Weight storage uses 671B. MLA KV uses kv_lora_rank=512.
Attention: MLA · layers: 61 · KV heads: 128. Architecture evidence: architecture.
Licence: DeepSeek Model License. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: safetensors, gguf. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2024-12-26 · source review recorded 2026-09-04.
Publisher model card and weightsMixtral 8x7B Instruct
46.7B total · Apache 2.0 · 32,768 context tokens
Architecture, licence and source
8 experts, 2 active per token. Weight memory uses 46.7B total, not 12.9B active.
Attention: GQA · layers: 32 · KV heads: 8. Architecture evidence: architecture.
Licence: Apache 2.0. Commercial-use claim: verified. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: gguf, safetensors. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2023-12-11 · source review recorded 2026-09-04.
Publisher model card and weightsMistral 7B Instruct v0.3
7.3B total · Apache 2.0 · 32,768 context tokens
Architecture, licence and source
No additional editorial note recorded.
Attention: GQA · layers: 32 · KV heads: 8. Architecture evidence: architecture.
Licence: Apache 2.0. Commercial-use claim: verified. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2024-05-22 · source review recorded 2026-09-04.
Publisher model card and weightsGemma 2 9B Instruct
9.2B total · Gemma Terms of Use · 8,192 context tokens
Architecture, licence and source
No additional editorial note recorded.
Attention: GQA · layers: 42 · KV heads: 8. Architecture evidence: architecture.
Licence: Gemma Terms of Use. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2024-06-27 · source review recorded 2026-09-04.
Publisher model card and weightsGemma 2 27B Instruct
27.2B total · Gemma Terms of Use · 8,192 context tokens
Architecture, licence and source
No additional editorial note recorded.
Attention: GQA · layers: 46 · KV heads: 16. Architecture evidence: architecture.
Licence: Gemma Terms of Use. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2024-06-27 · source review recorded 2026-09-04.
Publisher model card and weightsPhi-4
14.7B total · MIT · 16,384 context tokens
Architecture, licence and source
14.7B dense. MIT on the published Hugging Face card.
Attention: GQA · layers: 40 · KV heads: 10. Architecture evidence: architecture.
Licence: MIT. Commercial-use claim: verified. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.
Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.
Released 2024-12-12 · source review recorded 2026-09-04.
Publisher model card and weights