arunb74 commited on
Commit
b25f55b
·
verified ·
1 Parent(s): a20d50c

create README.md

Browse files
Files changed (1) hide show
  1. README.md +167 -0
README.md ADDED
@@ -0,0 +1,167 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ base_model: Nanbeige/Nanbeige4.2-3B
5
+ pipeline_tag: text-generation
6
+ tags:
7
+ - gguf
8
+ - llama.cpp
9
+ - nanbeige
10
+ - reasoning
11
+ - instruct
12
+ - local-llm
13
+ license: other
14
+ ---
15
+
16
+ # Nanbeige4.2-3B GGUF
17
+
18
+ This repository provides the **GGUF conversions** of the original **Nanbeige/Nanbeige4.2-3B** model. All credit for the model architecture and weights belongs to the original Nanbeige team.
19
+
20
+ > **Note:** This repository contains only GGUF conversions. The original Hugging Face model is **Nanbeige/Nanbeige4.2-3B**.
21
+
22
+ ---
23
+
24
+ # Base Model
25
+
26
+ **Original Hugging Face Model**
27
+
28
+ **Nanbeige/Nanbeige4.2-3B**
29
+
30
+ This repository does **not** modify or fine-tune the original model. It simply provides GGUF conversions for running the model locally with llama.cpp and other GGUF-compatible applications.
31
+
32
+ ---
33
+
34
+ # Available Files
35
+
36
+ This repository contains the following GGUF files:
37
+
38
+ | File | Description |
39
+ |------|-------------|
40
+ | **Nanbeige4.2-3B-BF16.gguf** | Full-precision BF16 GGUF. Highest quality but requires significantly more RAM/VRAM. |
41
+ | **Nanbeige4.2-3B-Q4_K_M.gguf** | 4-bit quantized GGUF. Recommended for most users due to its excellent balance of quality, speed, and memory usage. |
42
+
43
+ > **You only need to download ONE of these files.**
44
+
45
+ - Download **BF16** if you want the highest possible quality and have sufficient GPU memory.
46
+ - Download **Q4_K_M** if you want lower memory usage while maintaining excellent performance.
47
+
48
+ ---
49
+
50
+ # Important
51
+
52
+ At the time of publishing, support for the **Nanbeige** architecture has **not yet been merged into the main llama.cpp repository**.
53
+
54
+ Please use the official Nanbeige fork of llama.cpp:
55
+
56
+ https://github.com/Nanbeige/llama.cpp/tree/nanbeige42
57
+
58
+ If you build the upstream `ggml-org/llama.cpp`, you may encounter an error similar to:
59
+
60
+ ```text
61
+ unknown model architecture: 'nanbeige'
62
+ ```
63
+
64
+ The Nanbeige fork contains the required architecture support.
65
+
66
+ ---
67
+
68
+ # Building llama.cpp
69
+
70
+ Clone the repository:
71
+
72
+ ```bash
73
+ git clone --recursive -b nanbeige42 https://github.com/Nanbeige/llama.cpp.git
74
+ cd llama.cpp
75
+ ```
76
+
77
+ Build:
78
+
79
+ ```bash
80
+ cmake -B build
81
+ cmake --build build -j
82
+ ```
83
+
84
+ ---
85
+
86
+ # Running with llama-server
87
+
88
+ Example:
89
+
90
+ ```bash
91
+ ./build/bin/llama-server \
92
+ -m Nanbeige4.2-3B-Q4_K_M.gguf \
93
+ --host 0.0.0.0 \
94
+ --port 8080 \
95
+ -ngl 999 \
96
+ -c 65536
97
+ ```
98
+
99
+ After starting the server, the OpenAI-compatible API will be available at:
100
+
101
+ ```text
102
+ http://localhost:8080
103
+ ```
104
+
105
+ ---
106
+
107
+ # Using Hermes
108
+
109
+ Hermes works well with this model.
110
+
111
+ 1. Start `llama-server`.
112
+ 2. Open Hermes.
113
+ 3. Go to **Model Selection**.
114
+ 4. Choose **Custom URL**.
115
+ 5. Enter your llama-server endpoint, for example:
116
+
117
+ ```text
118
+ http://localhost:8080
119
+ ```
120
+
121
+ Hermes will communicate directly with your local Nanbeige model using the OpenAI-compatible API.
122
+
123
+ ---
124
+
125
+ # Using a Web UI
126
+
127
+ This model can also be used with web interfaces that support OpenAI-compatible endpoints, including:
128
+
129
+ - Open WebUI
130
+ - Hermes
131
+ - Any application compatible with the OpenAI Chat Completions API
132
+
133
+ Simply configure the application to connect to your running `llama-server` instance.
134
+
135
+ ---
136
+
137
+ # Recommended Model
138
+
139
+ For most users, the recommended file is:
140
+
141
+ **✅ Nanbeige4.2-3B-Q4_K_M.gguf**
142
+
143
+ It provides an excellent balance of:
144
+
145
+ - Quality
146
+ - Speed
147
+ - Memory usage
148
+
149
+ ---
150
+
151
+ # Credits
152
+
153
+ - **Original model:** Nanbeige/Nanbeige4.2-3B
154
+ - GGUF conversion provided by this repository.
155
+ - llama.cpp support is currently available through the Nanbeige `nanbeige42` branch.
156
+
157
+ All credit for the model architecture, tokenizer, training, and original model weights belongs entirely to the original Nanbeige team.
158
+
159
+ ---
160
+
161
+ # License
162
+
163
+ This repository distributes GGUF conversions of the original model.
164
+
165
+ The license for these GGUF files is **the same as the license of the original Nanbeige/Nanbeige4.2-3B project**.
166
+
167
+ Please refer to the original model repository for the official license terms, usage conditions, and any applicable restrictions.