ur-dad-matt commited on
Commit
be9b1c7
·
verified ·
1 Parent(s): e513abc

docs: pre-v1.8 banner + funnel to MLX lineup

Browse files
Files changed (1) hide show
  1. README.md +62 -46
README.md CHANGED
@@ -6,64 +6,80 @@ base_model: Qwen/Qwen2.5-32B-Instruct
6
  base_model_relation: quantized
7
  quantized_by: Outlier-Ai
8
  tags:
9
- - gguf
10
- - llama-cpp
11
- - llama.cpp
12
- - ollama
13
- - lm-studio
14
- - jan
15
- - quantized
16
- - 4bit
17
- - 4-bit
18
- - 5-bit
19
- - 8-bit
20
- - local-llm
21
- - on-device
22
- - cpu
23
- - edge-ai
24
- - offline
25
- - outlier
26
- - outlier-app
27
- - qwen2.5
28
- - qwen
29
- - text-generation
30
- - conversational
31
- - chat
32
- - instruct
33
- - windows
34
- - linux
35
- - mac
36
- - function-calling
37
  language:
38
- - en
39
- - zh
40
- - fr
41
- - es
42
- - pt
43
- - de
44
- - it
45
- - ru
46
- - ja
47
- - ko
48
- - ar
49
- - vi
50
- - th
51
- - nl
52
- - pl
53
  widget:
54
  - example_title: General knowledge
55
  messages:
56
  - role: user
57
- content: "Explain how mixture-of-experts models work in plain language, for a curious but non-technical reader."
 
58
  - example_title: Privacy rationale
59
  messages:
60
  - role: user
61
- content: "What are three practical reasons someone might want to run an LLM locally instead of in the cloud?"
 
62
  - example_title: Summarize
63
  messages:
64
  - role: user
65
- content: "Summarize the differences between 4-bit and 8-bit quantization in two short paragraphs."
 
66
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
67
 
68
  # Outlier Max 32B (GGUF)
69
 
 
6
  base_model_relation: quantized
7
  quantized_by: Outlier-Ai
8
  tags:
9
+ - gguf
10
+ - llama-cpp
11
+ - llama.cpp
12
+ - ollama
13
+ - lm-studio
14
+ - jan
15
+ - quantized
16
+ - 4bit
17
+ - 4-bit
18
+ - 5-bit
19
+ - 8-bit
20
+ - local-llm
21
+ - on-device
22
+ - cpu
23
+ - edge-ai
24
+ - offline
25
+ - outlier
26
+ - outlier-app
27
+ - qwen2.5
28
+ - qwen
29
+ - text-generation
30
+ - conversational
31
+ - chat
32
+ - instruct
33
+ - windows
34
+ - linux
35
+ - mac
36
+ - function-calling
37
  language:
38
+ - en
39
+ - zh
40
+ - fr
41
+ - es
42
+ - pt
43
+ - de
44
+ - it
45
+ - ru
46
+ - ja
47
+ - ko
48
+ - ar
49
+ - vi
50
+ - th
51
+ - nl
52
+ - pl
53
  widget:
54
  - example_title: General knowledge
55
  messages:
56
  - role: user
57
+ content: Explain how mixture-of-experts models work in plain language, for a curious
58
+ but non-technical reader.
59
  - example_title: Privacy rationale
60
  messages:
61
  - role: user
62
+ content: What are three practical reasons someone might want to run an LLM locally
63
+ instead of in the cloud?
64
  - example_title: Summarize
65
  messages:
66
  - role: user
67
+ content: Summarize the differences between 4-bit and 8-bit quantization in two
68
+ short paragraphs.
69
  ---
70
+ > # ⚠️ Pre-v1.8 build — see [Outlier-Ai org](https://huggingface.co/Outlier-Ai) for current MLX 4-bit lineup
71
+ >
72
+ > This is a **Max (deferred to v1.9+ — see Plus tier on Outlier-Ai org) GGUF** build from a pre-v1.8 era. The current Outlier app (v1.8+) ships **MLX 4-bit** versions optimized for Apple Silicon — much faster than GGUF on Mac.
73
+ >
74
+ > | Want | Use |
75
+ > |---|---|
76
+ > | **Run on macOS** (any M-series Mac) | [Outlier-Ai (MLX 4-bit lineup)](https://huggingface.co/Outlier-Ai) |
77
+ > | **Run on Windows / Linux** | This GGUF still works with [llama.cpp](https://github.com/ggerganov/llama.cpp), [Ollama](https://ollama.ai), [LM Studio](https://lmstudio.ai), [Jan](https://jan.ai) |
78
+ >
79
+ > [📥 Download the Outlier desktop app](https://outlier.host)
80
+
81
+ ---
82
+
83
 
84
  # Outlier Max 32B (GGUF)
85