Small enough for an ordinary Mac, and each one actually tested in Korean before release. Chat, OCR, speech and embeddings.
AI & ML interests
On-device AI, GGUF quantization, Apple Silicon, macOS automation
Recent Activity
View all activity
Qwen's newest 27B: vision, 262K context, Korean-verified. Start with IQ4_XS β we measured every tier and Q3_K_M is smaller yet slower. Needs 24GB+.
NVIDIA Nemotron 3 family β NemotronH architecture combining Mamba state-space + standard attention. Mac-runnable, BatiAI-quantized + signed.
Latest Qwen 3.6 series with native tool calling, thinking mode, and Vision-Language. Best balance for 48-128GB Macs.
Qwen 3.5 dense and MoE quantizations. Reliable tool calling and JSON generation.
Korean output, tool calling and load verified on every tier before publishing. Failures documented on the cards, not omitted.
batisay(STT, 무μμ λ§νλ) + batispeak(νμλΆλ¦¬, λκ° λ§νλ) = ν΅νΒ·νμ νμλ³ μ μ¬. 16GB Mac on-device.
Largest open-weight LLMs, BatiAI-quantized. Mac-runnable from M4 Max 128GB to Mac Studio M3 Ultra 512GB.
Gemma 4 quantizations from Google's official weights. Best entry for 16GB Mac mini M4 (E4B Q4 = 57 t/s).
-
batiai/gemma-4-E2B-it-GGUF
Text Generation β’ 5B β’ Updated β’ 466 β’ 2 -
batiai/gemma-4-E4B-it-GGUF
Text Generation β’ 8B β’ Updated β’ 712 β’ 5 -
batiai/Gemma-4-26B-A4B-it-GGUF
Text Generation β’ 25B β’ Updated β’ 923 β’ 3 -
batiai/gemma-4-31B-it-GGUF
Text Generation β’ 31B β’ Updated β’ 281
Complete Mac-first on-device RAG stack β chat LLM + reranker + text/VL embedder, direct from BF16, BatiAI-signed. For BatiFlow.
-
batiai/Qwen3-Embedding-4B-GGUF
Sentence Similarity β’ 4B β’ Updated β’ 217 -
batiai/Qwen3-Embedding-0.6B-GGUF
Sentence Similarity β’ 0.6B β’ Updated β’ 241 β’ 1 -
batiai/Qwen3-VL-Embedding-8B-GGUF
Sentence Similarity β’ 8B β’ Updated β’ 945 β’ 7 -
batiai/Qwen3-VL-Embedding-2B-GGUF
Sentence Similarity β’ 2B β’ Updated β’ 118
Small enough for an ordinary Mac, and each one actually tested in Korean before release. Chat, OCR, speech and embeddings.
Korean output, tool calling and load verified on every tier before publishing. Failures documented on the cards, not omitted.
Qwen's newest 27B: vision, 262K context, Korean-verified. Start with IQ4_XS β we measured every tier and Q3_K_M is smaller yet slower. Needs 24GB+.
batisay(STT, 무μμ λ§νλ) + batispeak(νμλΆλ¦¬, λκ° λ§νλ) = ν΅νΒ·νμ νμλ³ μ μ¬. 16GB Mac on-device.
NVIDIA Nemotron 3 family β NemotronH architecture combining Mamba state-space + standard attention. Mac-runnable, BatiAI-quantized + signed.
Largest open-weight LLMs, BatiAI-quantized. Mac-runnable from M4 Max 128GB to Mac Studio M3 Ultra 512GB.
Latest Qwen 3.6 series with native tool calling, thinking mode, and Vision-Language. Best balance for 48-128GB Macs.
Gemma 4 quantizations from Google's official weights. Best entry for 16GB Mac mini M4 (E4B Q4 = 57 t/s).
-
batiai/gemma-4-E2B-it-GGUF
Text Generation β’ 5B β’ Updated β’ 466 β’ 2 -
batiai/gemma-4-E4B-it-GGUF
Text Generation β’ 8B β’ Updated β’ 712 β’ 5 -
batiai/Gemma-4-26B-A4B-it-GGUF
Text Generation β’ 25B β’ Updated β’ 923 β’ 3 -
batiai/gemma-4-31B-it-GGUF
Text Generation β’ 31B β’ Updated β’ 281
Qwen 3.5 dense and MoE quantizations. Reliable tool calling and JSON generation.
Complete Mac-first on-device RAG stack β chat LLM + reranker + text/VL embedder, direct from BF16, BatiAI-signed. For BatiFlow.
-
batiai/Qwen3-Embedding-4B-GGUF
Sentence Similarity β’ 4B β’ Updated β’ 217 -
batiai/Qwen3-Embedding-0.6B-GGUF
Sentence Similarity β’ 0.6B β’ Updated β’ 241 β’ 1 -
batiai/Qwen3-VL-Embedding-8B-GGUF
Sentence Similarity β’ 8B β’ Updated β’ 945 β’ 7 -
batiai/Qwen3-VL-Embedding-2B-GGUF
Sentence Similarity β’ 2B β’ Updated β’ 118