pico-type πŸ”

A tiny byte-level multi-head content classifier β€” ~1.5M params, ~9MB single-file ONNX (FP32), ~18ms CPU inference.

Classifies any content from raw bytes: coarse type Β· modality Β· subtype Β· code language Β· text language Β· file MIME Β· risk flags

License Python PyPI ONNX CI HuggingFace Space HuggingFace Model


✨ Features

  • No tokenizer β€” operates directly on raw UTF-8 bytes (supports all languages, no preprocessing)
  • 7 heads, one forward pass β€” coarse type, modality, subtype, code language, text language, file MIME, risk flags
  • 4 Matryoshka tiers β€” tiny (16d) β†’ small (64d) β†’ base (192d) β†’ pro (576d) β€” same trunk, accuracy scales with dim
  • ~9MB single-file ONNX (FP32) β€” deploy on edge devices, serverless, browser (WebAssembly/ONNX Runtime Web)
  • ~18ms inference on CPU via ONNX Runtime
  • CLI, Python API, Gradio Space, MCP server β€” ready to use

πŸ“Š Evaluation

Overall Accuracy (v2 β€” trained on real data)

Head Classes Accuracy Dataset
coarse 12 100% Synthetic eval
modality 8 100% Synthetic eval
subtype 24 93.8% Synthetic eval
code_lang 62 60.3% The Heap β€” 24 real-world langs, 1,200 samples
text_lang 30 98.3% Wikipedia β€” 30 langs, 1,500 samples
file_mime 90 100% Synthetic eval
risk (mAP) 6 100% Synthetic eval

v0.1 baseline (synthetic-only): code_lang 3%, text_lang 19%. Real-data training in v2 improves code by 57pp and text by 79pp.

Code Language β€” Per-Language Accuracy

Excellent (90%+) Good (70–89%) Needs Work (<50%)
cpp 96%, dart 98%, erlang 98%, rust 98%, r 94%, swift 92%, python 88%, lua 88% go 86%, ruby 86%, ocaml 84%, php 78%, csharp 76%, java 76%, kotlin 76%, c 62% perl 50%, haskell 24%, scala 4%, javascript 2%, clojure 0%, elixir 0%, julia 0%, sql 0%

Note: Low-accuracy languages have fewer real training samples. More data will improve them.

πŸš€ Quick Start

Install

pip install picotype

CLI

# Classify from stdin
echo "def hello(name):\n    return f'Hi {name}'" | picotype --pretty

# Classify a file
picotype --file document.txt

# Classify clipboard content
picotype --clip

# All 4 tiers available
echo "..." | picotype --tier pro

Python API

from picotype import load_onnx_model, run_onnx

session = load_onnx_model("base")
result = run_onnx(session, "def hello(): pass")
print(result)
# {
#   "coarse": "code",
#   "code_language": "python",
#   "modality": "textual",
#   "confidence": 0.98,
#   ...
# }

MCP Server (for Claude Desktop, Cursor, etc.)

pip install picotype
PICOTYPE_MODEL_DIR=./checkpoints python -m model.pico_type.mcp_server

Then add to your MCP config:

{
  "mcpServers": {
    "pico-type": {
      "command": "python",
      "args": ["-m", "model.pico_type.mcp_server"],
      "env": { "PICOTYPE_MODEL_DIR": "./checkpoints" }
    }
  }
}

Gradio Web UI

Try it live: huggingface.co/spaces/eulogik/pico-type

πŸ— Architecture

Bytes ─▢ ByteEmbed(256β†’96d) ─▢ 3Γ—Conv1D(k=3,5,7) ─▢ 2Γ—BiAttention(RoPE) ─▢ Pool ─▢ 7Γ—Matryoshka Heads
Component Detail
ByteEmbed Lookup-free embedding β€” each byte value (0–255) maps to a learned 96-dim vector
Conv1D 3 parallel depthwise convolutions (kernel widths 3, 5, 7) with residual + layer norm
BiAttention Bidirectional self-attention with Rotary Position Embeddings (RoPE), 4 heads
Pool Mean + max + std deviation concatenation β†’ fixed-size representation
Heads Matryoshka-style: slice pool dim to 16/64/192/576, project to 7 linear classifiers

Total parameters: 1.43M (tiny) / 1.45M (small) / 1.48M (base) / 1.56M (pro)

πŸ”§ Model Tiers

Tier Dim Params ONNX Size Accuracy Multiplier
tiny 16 1.43M 9.09 MB 0.65Γ—
small 64 1.45M 9.13 MB 0.82Γ—
base 192 1.48M 9.25 MB 1.0Γ— (reference)
pro 576 1.56M 9.61 MB 1.05Γ—

ONNX sizes are single-file FP32 exports (graph-only files are 203–206 KB).

All tiers share the same backbone; only the final linear projection layers differ. Higher-tier models use more dimensions for finer-grained classification.

πŸ§ͺ Classification Heads

Head Classes What It Detects
coarse 12 text, code, link, image, file, config, markup, data, error, secret, archive, binary
modality 8 textual, binary_image, binary_archive, binary_executable, binary_document, etc.
subtype 24 json, yaml, toml, csv, html, markdown, sql, log, dockerfile, makefile, etc.
code_lang 62 python, javascript, typescript, java, c, cpp, go, rust, ruby, php, swift, kotlin, and 50 more
text_lang 30 en, es, fr, de, it, pt, nl, ru, zh, ja, ko, vi, th, id, and 15 more
file_mime 90 application/json, image/png, video/mp4, font/ttf, application/wasm, and 84 more
risk 6 api_key, jwt, password, email, phone, ssh_key

🌐 Deployment

Platform Link Notes
HuggingFace Space eulogik/pico-type Gradio web UI, no GPU needed
HuggingFace Model eulogik/pico-type ONNX models + export metadata
GitHub eulogik/pico-type Source code, training, paper
PyPI pip install picotype Python package
ONNX Runtime Use with onnxruntime.js Browser/Node.js deployment

πŸ“š Resources

  • Paper β€” Architecture, training, and evaluation details
  • Model Card β€” Detailed architecture and training configuration
  • Walkthrough β€” Development log and decisions
  • Architecture Plan β€” Original design document

πŸ“„ License

Apache 2.0


Built with PyTorch Β· ONNX Β· Gradio Β· HuggingFace
Downloads last month
162
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including eulogik/pico-type