Open weights

Locally runnable models

Explore downloadable models by architecture, parameter count, published context, licence and recorded weight formats. Specifications do not establish whether a model will fit or run well on a particular computer.

Catalogue review recorded · individual claims have separate datesStructured datasetSourceMethodology

A context ceiling is not a retrieval-quality result. For mixture-of-experts models, active parameters describe the subset used during a token step; they are not the total weights that must be stored. No model-fit verdict is inferred from these specifications.

Llama 3.1 8B Instruct

8.0B total · Llama 3.1 Community License · 131,072 context tokens

Architecture, licence and source

GQA 8 KV heads. Weight memory uses 8.03B total parameters.

Attention: GQA · layers: 32 · KV heads: 8. Architecture evidence: architecture.

Licence: Llama 3.1 Community License. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2024-07-23 · source review recorded 2026-09-04.

Publisher model card and weights

Llama 3.1 70B Instruct

70.6B total · Llama 3.1 Community License · 131,072 context tokens

Architecture, licence and source

No additional editorial note recorded.

Attention: GQA · layers: 80 · KV heads: 8. Architecture evidence: architecture.

Licence: Llama 3.1 Community License. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2024-07-23 · source review recorded 2026-09-04.

Publisher model card and weights

Llama 3.3 70B Instruct

70.6B total · Llama 3.3 Community License · 131,072 context tokens

Architecture, licence and source

Same 70B GQA layout as Llama 3.1 70B; instruction weights differ.

Attention: GQA · layers: 80 · KV heads: 8. Architecture evidence: architecture.

Licence: Llama 3.3 Community License. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2024-12-06 · source review recorded 2026-09-04.

Publisher model card and weights

Qwen2.5 7B Instruct

7.6B total · Apache 2.0 · 131,072 context tokens

Architecture, licence and source

Native 32K; 128K via YaRN as documented on the model card.

Attention: GQA · layers: 28 · KV heads: 4. Architecture evidence: architecture.

Licence: Apache 2.0. Commercial-use claim: verified. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2024-09-19 · source review recorded 2026-09-04.

Publisher model card and weights

Qwen2.5 32B Instruct

32.5B total · Apache 2.0 · 131,072 context tokens

Architecture, licence and source

No additional editorial note recorded.

Attention: GQA · layers: 64 · KV heads: 8. Architecture evidence: architecture.

Licence: Apache 2.0. Commercial-use claim: verified. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2024-09-19 · source review recorded 2026-09-04.

Publisher model card and weights

Qwen2.5 72B Instruct

72.7B total · Qwen License · 131,072 context tokens

Architecture, licence and source

Tongyi Qianwen license — not Apache 2.0. Commercial use follows the published license.

Attention: GQA · layers: 80 · KV heads: 8. Architecture evidence: architecture.

Licence: Qwen License. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2024-09-19 · source review recorded 2026-09-04.

Publisher model card and weights

Qwen2.5-Coder 32B Instruct

32.5B total · Apache 2.0 · 131,072 context tokens

Architecture, licence and source

No additional editorial note recorded.

Attention: GQA · layers: 64 · KV heads: 8. Architecture evidence: architecture.

Licence: Apache 2.0. Commercial-use claim: verified. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2024-11-11 · source review recorded 2026-09-04.

Publisher model card and weights

DeepSeek-V3

671B total · DeepSeek Model License · 131,072 context tokens

Architecture, licence and source

MoE 671B total / 37B active. Weight storage uses 671B. MLA KV uses kv_lora_rank=512.

Attention: MLA · layers: 61 · KV heads: 128. Architecture evidence: architecture.

Licence: DeepSeek Model License. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: safetensors, gguf. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2024-12-26 · source review recorded 2026-09-04.

Publisher model card and weights

Mixtral 8x7B Instruct

46.7B total · Apache 2.0 · 32,768 context tokens

Architecture, licence and source

8 experts, 2 active per token. Weight memory uses 46.7B total, not 12.9B active.

Attention: GQA · layers: 32 · KV heads: 8. Architecture evidence: architecture.

Licence: Apache 2.0. Commercial-use claim: verified. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: gguf, safetensors. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2023-12-11 · source review recorded 2026-09-04.

Publisher model card and weights

Mistral 7B Instruct v0.3

7.3B total · Apache 2.0 · 32,768 context tokens

Architecture, licence and source

No additional editorial note recorded.

Attention: GQA · layers: 32 · KV heads: 8. Architecture evidence: architecture.

Licence: Apache 2.0. Commercial-use claim: verified. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2024-05-22 · source review recorded 2026-09-04.

Publisher model card and weights

Gemma 2 9B Instruct

9.2B total · Gemma Terms of Use · 8,192 context tokens

Architecture, licence and source

No additional editorial note recorded.

Attention: GQA · layers: 42 · KV heads: 8. Architecture evidence: architecture.

Licence: Gemma Terms of Use. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2024-06-27 · source review recorded 2026-09-04.

Publisher model card and weights

Gemma 2 27B Instruct

27.2B total · Gemma Terms of Use · 8,192 context tokens

Architecture, licence and source

No additional editorial note recorded.

Attention: GQA · layers: 46 · KV heads: 16. Architecture evidence: architecture.

Licence: Gemma Terms of Use. Commercial-use claim: official docs. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2024-06-27 · source review recorded 2026-09-04.

Publisher model card and weights

Phi-4

14.7B total · MIT · 16,384 context tokens

Architecture, licence and source

14.7B dense. MIT on the published Hugging Face card.

Attention: GQA · layers: 40 · KV heads: 10. Architecture evidence: architecture.

Licence: MIT. Commercial-use claim: verified. Check the actual licence and acceptable-use restrictions; downloadable weights do not themselves grant unrestricted use.

Formats recorded: gguf, safetensors, mlx. A runtime accepting a format does not establish support for every architecture or quantization.

Released 2024-12-12 · source review recorded 2026-09-04.

Publisher model card and weights