JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
Abstract
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state. Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift toward agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time. To address this gap, we introduce JarvisHub, a canvas-native creative agent harness for long-horizon multimodal creation. JarvisHub treats an editable canvas as the user workspace, the agent's external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use toward sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.
Community
Our new project, JarvisHub, is now live and open source! đ„ł
Unlike one-shot prompt-based generation tools, linear chatbot agents, or node-based workflows that require manual setup, JarvisHub is a Canvas-Native Agent Harness for long-horizon multimodal creation. It turns an editable canvas into a shared project state where users and agents can collaborate seamlessly.
â
Canvas as Memory: Prompts, reference materials, candidate versions, generated outputs, and feedback all remain on the canvasâallowing the agent to truly âseeâ the entire project and continuously move it forward.
â
Inspectable and Recoverable: The generation process, dependencies, and revision history are clearly visible. When something goes wrong, there is no need to start overâsimply fix the relevant node.
â
Long-Horizon Multimodal Creation: JarvisHub natively supports images, videos, websites, presentations, and more, enabling agents to move beyond âgenerating an outputâ toward âcompleting an entire project.â đ
JarvisHub currently showcases three representative use cases: narrative media generation, interactive web development, and presentation creation.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing (2026)
- VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation (2026)
- COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows (2026)
- CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration (2026)
- Plover: Steering GUI Agents through Plan-Centric Interaction (2026)
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources (2026)
- MemoGen: Can Past Experience Improve Future Text-to-Image Generation? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.23588 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper