pico-type π
A tiny byte-level multi-head content classifier β ~1.5M params, ~9MB single-file ONNX (FP32), ~18ms CPU inference.
Classifies any content from raw bytes: coarse type Β· modality Β· subtype Β· code language Β· text language Β· file MIME Β· risk flags
β¨ Features
- No tokenizer β operates directly on raw UTF-8 bytes (supports all languages, no preprocessing)
- 7 heads, one forward pass β coarse type, modality, subtype, code language, text language, file MIME, risk flags
- 4 Matryoshka tiers β tiny (16d) β small (64d) β base (192d) β pro (576d) β same trunk, accuracy scales with dim
- ~9MB single-file ONNX (FP32) β deploy on edge devices, serverless, browser (WebAssembly/ONNX Runtime Web)
- ~18ms inference on CPU via ONNX Runtime
- CLI, Python API, Gradio Space, MCP server β ready to use
π Evaluation
Overall Accuracy (v2 β trained on real data)
| Head | Classes | Accuracy | Dataset |
|---|---|---|---|
| coarse | 12 | 100% | Synthetic eval |
| modality | 8 | 100% | Synthetic eval |
| subtype | 24 | 93.8% | Synthetic eval |
| code_lang | 62 | 60.3% | The Heap β 24 real-world langs, 1,200 samples |
| text_lang | 30 | 98.3% | Wikipedia β 30 langs, 1,500 samples |
| file_mime | 90 | 100% | Synthetic eval |
| risk (mAP) | 6 | 100% | Synthetic eval |
v0.1 baseline (synthetic-only): code_lang 3%, text_lang 19%. Real-data training in v2 improves code by 57pp and text by 79pp.
Code Language β Per-Language Accuracy
| Excellent (90%+) | Good (70β89%) | Needs Work (<50%) |
|---|---|---|
| cpp 96%, dart 98%, erlang 98%, rust 98%, r 94%, swift 92%, python 88%, lua 88% | go 86%, ruby 86%, ocaml 84%, php 78%, csharp 76%, java 76%, kotlin 76%, c 62% | perl 50%, haskell 24%, scala 4%, javascript 2%, clojure 0%, elixir 0%, julia 0%, sql 0% |
Note: Low-accuracy languages have fewer real training samples. More data will improve them.
π Quick Start
Install
pip install picotype
CLI
# Classify from stdin
echo "def hello(name):\n return f'Hi {name}'" | picotype --pretty
# Classify a file
picotype --file document.txt
# Classify clipboard content
picotype --clip
# All 4 tiers available
echo "..." | picotype --tier pro
Python API
from picotype import load_onnx_model, run_onnx
session = load_onnx_model("base")
result = run_onnx(session, "def hello(): pass")
print(result)
# {
# "coarse": "code",
# "code_language": "python",
# "modality": "textual",
# "confidence": 0.98,
# ...
# }
MCP Server (for Claude Desktop, Cursor, etc.)
pip install picotype
PICOTYPE_MODEL_DIR=./checkpoints python -m model.pico_type.mcp_server
Then add to your MCP config:
{
"mcpServers": {
"pico-type": {
"command": "python",
"args": ["-m", "model.pico_type.mcp_server"],
"env": { "PICOTYPE_MODEL_DIR": "./checkpoints" }
}
}
}
Gradio Web UI
Try it live: huggingface.co/spaces/eulogik/pico-type
π Architecture
Bytes ββΆ ByteEmbed(256β96d) ββΆ 3ΓConv1D(k=3,5,7) ββΆ 2ΓBiAttention(RoPE) ββΆ Pool ββΆ 7ΓMatryoshka Heads
| Component | Detail |
|---|---|
| ByteEmbed | Lookup-free embedding β each byte value (0β255) maps to a learned 96-dim vector |
| Conv1D | 3 parallel depthwise convolutions (kernel widths 3, 5, 7) with residual + layer norm |
| BiAttention | Bidirectional self-attention with Rotary Position Embeddings (RoPE), 4 heads |
| Pool | Mean + max + std deviation concatenation β fixed-size representation |
| Heads | Matryoshka-style: slice pool dim to 16/64/192/576, project to 7 linear classifiers |
Total parameters: 1.43M (tiny) / 1.45M (small) / 1.48M (base) / 1.56M (pro)
π§ Model Tiers
| Tier | Dim | Params | ONNX Size | Accuracy Multiplier |
|---|---|---|---|---|
| tiny | 16 | 1.43M | 9.09 MB | 0.65Γ |
| small | 64 | 1.45M | 9.13 MB | 0.82Γ |
| base | 192 | 1.48M | 9.25 MB | 1.0Γ (reference) |
| pro | 576 | 1.56M | 9.61 MB | 1.05Γ |
ONNX sizes are single-file FP32 exports (graph-only files are 203β206 KB).
All tiers share the same backbone; only the final linear projection layers differ. Higher-tier models use more dimensions for finer-grained classification.
π§ͺ Classification Heads
| Head | Classes | What It Detects |
|---|---|---|
| coarse | 12 | text, code, link, image, file, config, markup, data, error, secret, archive, binary |
| modality | 8 | textual, binary_image, binary_archive, binary_executable, binary_document, etc. |
| subtype | 24 | json, yaml, toml, csv, html, markdown, sql, log, dockerfile, makefile, etc. |
| code_lang | 62 | python, javascript, typescript, java, c, cpp, go, rust, ruby, php, swift, kotlin, and 50 more |
| text_lang | 30 | en, es, fr, de, it, pt, nl, ru, zh, ja, ko, vi, th, id, and 15 more |
| file_mime | 90 | application/json, image/png, video/mp4, font/ttf, application/wasm, and 84 more |
| risk | 6 | api_key, jwt, password, email, phone, ssh_key |
π Deployment
| Platform | Link | Notes |
|---|---|---|
| HuggingFace Space | eulogik/pico-type | Gradio web UI, no GPU needed |
| HuggingFace Model | eulogik/pico-type | ONNX models + export metadata |
| GitHub | eulogik/pico-type | Source code, training, paper |
| PyPI | pip install picotype |
Python package |
| ONNX Runtime | Use with onnxruntime.js | Browser/Node.js deployment |
π Resources
- Paper β Architecture, training, and evaluation details
- Model Card β Detailed architecture and training configuration
- Walkthrough β Development log and decisions
- Architecture Plan β Original design document
π License
Apache 2.0
- Downloads last month
- 162