Image-Text-to-Text
Transformers
Safetensors
Rust
English
muse_glimmer
esper
esper-4
valiant
valiant-labs
meta
facebook
muse-glimmer
muse
glimmer
muse-glimmer-30b
30b
reasoning
code
code-instruct
python
typescript
javascript
java
c++
c
c#
go
haskell
dev-ops
jenkins
terraform
ansible
docker
kubernetes
helm
grafana
prometheus
shell
bash
azure
aws
gcp
cloud
scripting
powershell
problem-solving
architect
engineer
developer
creative
analytical
expert
rationality
conversational
chat
instruct
Instructions to use ValiantLabs/Muse-Glimmer-30B-Esper4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ValiantLabs/Muse-Glimmer-30B-Esper4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="ValiantLabs/Muse-Glimmer-30B-Esper4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ValiantLabs/Muse-Glimmer-30B-Esper4") model = AutoModelForMultimodalLM.from_pretrained("ValiantLabs/Muse-Glimmer-30B-Esper4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ValiantLabs/Muse-Glimmer-30B-Esper4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ValiantLabs/Muse-Glimmer-30B-Esper4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ValiantLabs/Muse-Glimmer-30B-Esper4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ValiantLabs/Muse-Glimmer-30B-Esper4
- SGLang
How to use ValiantLabs/Muse-Glimmer-30B-Esper4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ValiantLabs/Muse-Glimmer-30B-Esper4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ValiantLabs/Muse-Glimmer-30B-Esper4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ValiantLabs/Muse-Glimmer-30B-Esper4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ValiantLabs/Muse-Glimmer-30B-Esper4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use ValiantLabs/Muse-Glimmer-30B-Esper4 with Docker Model Runner:
docker model run hf.co/ValiantLabs/Muse-Glimmer-30B-Esper4
File size: 4,860 Bytes
1f5f251 60d66ba 1f5f251 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 | ---
language:
- en
library_name: transformers
pipeline_tag: image-text-to-text
tags:
- esper
- esper-4
- valiant
- valiant-labs
- meta
- facebook
- muse-glimmer
- muse
- glimmer
- muse-glimmer-30b
- 30b
- reasoning
- code
- code-instruct
- python
- typescript
- javascript
- java
- c++
- c
- c#
- rust
- go
- haskell
- dev-ops
- jenkins
- terraform
- ansible
- docker
- jenkins
- kubernetes
- helm
- grafana
- prometheus
- shell
- bash
- azure
- aws
- gcp
- cloud
- scripting
- powershell
- problem-solving
- architect
- engineer
- developer
- creative
- analytical
- expert
- rationality
- conversational
- chat
- instruct
base_model: meta-models/Muse-Glimmer-30B
datasets:
- sequelbox/Mitakihara2-DeepSeek-V4-Pro
- sequelbox/Tachibana4-DeepSeek-V4-Pro
- sequelbox/Titanium4-DeepSeek-V4-Pro
license: apache-2.0
---
**[Support our open-source dataset and model releases!](https://huggingface.co/spaces/sequelbox/SupportOpenSource)**

Esper 4: [gemma-4-12B](https://huggingface.co/ValiantLabs/gemma-4-12B-it-Esper4), [Qwen3.6-27B](https://huggingface.co/ValiantLabs/Qwen3.6-27B-Esper4), [Muse-Glimmer-30B](https://huggingface.co/ValiantLabs/Muse-Glimmer-30B-Esper4)
Esper 4 is an agentic coding, architecture, DevOps, and MLOps specialist built on Muse Glimmer 30B!
- Your dedicated DevOps expert: Esper 4 maximizes DevOps and architecture helpfulness, powered by [high-difficulty DevOps and architecture data](https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro) generated with DeepSeek-V4-Pro!
- Improved coding performance: [challenging agentic coding queries](https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro) allow Esper 4 to tackle harder coding tasks!
- AI to build AI: our [high-difficulty AI coding and expertise data](https://huggingface.co/datasets/sequelbox/Mitakihara2-DeepSeek-V4-Pro) boosts Esper 4 for AI development, research, deployment, interpretability, operation and experimentation!
- Small model sizes allow running on local desktop and mobile, plus super-fast server inference!
## Prompting Guide
Esper 4 uses the [Muse Glimmer](https://huggingface.co/meta-models/Muse-Glimmer-30B) prompt format.
Use Esper 4 with your agentic framework of choice or as a stand-alone chat and code assistant.
Example inference script to get started:
```python
from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "ValiantLabs/Muse-Glimmer-30B-Esper4"
# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(MODEL_ID, dtype="auto", device_map="auto")
# Prompt
prompt = "Implement CQRS for network appliance config management.\n\nRequirements:\n- Write side: 200 commands/sec, 4 command handlers, SQLite with custom journaling\n- Read side: 1000 queries/sec, 3 read projections in shared memory segments\n- Eventual consistency window: 100ms max\n- Handle atomic swap of projection memory for rebuilds\n- Binary configuration format versioning for schema evolution\n- Framework: libevent with custom protocol parser\n\nConstraints:\n- Manual memory management only, no garbage collection\n- Lock-free data structures where possible\n- Shared memory projections must survive process restarts\n- Command handlers must be thread-safe with 4 worker threads\n- Projection rebuild must not block queries\n- Binary format must support forward/backward compatibility\n- Error handling for corrupted journal recovery\n- Memory-mapped I/O for shared segments\n- Zero-copy where possible for performance\n\nDeliverables:\n1. Command processing pipeline with journaling\n2. Projection engine with shared memory management\n3. Query dispatcher with read-your-writes consistency\n4. Schema evolution system with versioned binary format\n5. Integration with libevent for network I/O\n6. Stress test showing 200 cmd/s + 1000 q/s sustained\n\nAssume x86_64 Linux, pthreads, atomic operations. No high-level frameworks."
messages = [
{"role": "user", "content": prompt},
]
# Process input
inputs = processor.apply_chat_template(
messages,
tokenize=True,
return_dict=True,
return_tensors="pt",
add_generation_prompt=True,
reasoning_strength="high",
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
# Generate output
outputs = model.generate(**inputs)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
print(response)
```

Esper 4 is created by [Valiant Labs.](http://valiantlabs.ca/)
[Check out our HuggingFace page to see all of our models!](https://huggingface.co/ValiantLabs)
We care about open source. For everyone to use.
|