True2456/nemotron-3-super-120b-cybersecurity-theory-lora-mlx Text Generation • Updated 15 days ago • 2
Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms Paper • 2607.07769 • Published 24 days ago • 10
Running Featured 405 Bonsai 27B WebGPU Kernels 🌳 405 Run a 1-bit 27B LLM locally in your browser on WebGPU
Running 205 The ultimate guide to RL environments: building and scaling them in the LLM era 📝 205 Building and scaling RL environments for LLM training