
Security-first autonomous coding agent
A self-contained autonomous coding agent — with a cinematic soul. A complete dev environment with a Deep Agents orchestrator, sandboxed skills, and persistent Markdown memory.
System Heartbeat
The Personality
madz is not a blank slate. It is built on the cinematic soul of Mads Mikkelsen — five characters, selected by the task at hand. It is not decoration; it is the difference between a tool that answers and a teammate that engages.
Code review, security audit, architectural critique, critical analysis
Debugging, tracing, mathematical problems, error analysis
Building, fixing, designing systems, scaffolding
Brainstorming, exploring, when you are stuck, creative problem-solving
Calm decisiveness under pressure, incident response, high-stakes decisions
The default is Hannibal — the one it returns to. The others are situational, selected by the task at hand. It speaks with a measured cadence, dry humor, and a genuine point of view.
Why Madz
Most coding agents are a thin layer over a model. madz is the opposite — a complete development environment with an agent at its center. That distinction is the whole point.
A full toolchain ships in the container — Node.js 26, Python 3, Ruby, Go, Java (OpenJDK 21), Rust, Terraform, tflint, Chromium, ripgrep, git, the GitHub CLI, and npm, yarn, and pnpm. A working dev box, not a chat window.
It remembers your context, adapts to your energy, and works with you rather than at you. Long sessions feel collaborative instead of transactional.
Skills run in isolated child processes with time limits, memory caps, and allowlists for filesystem paths and outbound URLs. Every tool is gated by explicit permissions, and sensitive fields are redacted from telemetry.
A Deep Agents orchestrator routes work to a family of specialized subagents. It can read your codebase, find bugs, write fixes, run tests, and open PRs without you driving every step.
Memory, context, sessions, schedules, and skill definitions are plain Markdown under version control. git log is your audit trail — essential where you must prove what happened and why.
Data stays in your environment. Nothing leaves the container unless you explicitly configure it to. Secrets are loaded only from environment variables and never logged or hardcoded.
How It Works
The orchestrator routes tasks to specialized subagents, each with a focused system prompt and a curated tool set. Tools are gated by permissions and sandbox constraints, skills execute in isolated child processes, and everything is persisted as Markdown and indexed for semantic search.
Deep Agents
Orchestrator
primary agent + 12 specialized subagents
Sandboxed skills
isolated child process
LLM provider
OpenAI-compatible
Built-in tools
30+ · deepagents filesystem
Specialized Subagents
Features
Not a promise that the agent is harmless, but a guarantee that untrusted code is constrained — through permission gating, sandboxed execution, and auditable Markdown state.
First launch collects your profile — attractor, expertise, dev tools, communication style — into memory/context/profile.md, loaded into every session.
Configurable provider dispatch with rate limiting and context-window trimming. Supports OpenAI-compatible APIs.
A cache-aside LRU sits between agent and provider, keyed on thread ID and a SHA-256 of the message. Fail-open, and skipped whenever tools ran.
Deep Agents orchestrates a primary agent with a family of specialized subagents, each with a focused prompt and curated tool set.
30+ tools for shell processes, spreadsheets, PDFs, web search and extraction, webhooks, YAML, and speech — plus semantic code search.
Auto-discovers Agent Skills spec skills from a skills/ directory. Each carries SKILL.md frontmatter and optional executable scripts.
Tools register only when their permissions are enabled — filesystem read, write, and exec, process spawn, and network outbound.
Skills run in isolated spawned child processes with time limits, memory caps, path allowlists, and blocked URL schemes.
Optional OpenTelemetry integration with console, OTLP HTTP, or OTLP gRPC exporters, probability sampling, and automatic redaction.
Semantic code search over local embeddings and SQLite KNN retrieval — find code by meaning rather than exact keywords.
Declare recurring jobs in config.yaml. In-process or system crontab delegation, with max-concurrency control to prevent overlap.
A triple-layer architecture of canonical, ephemeral, and reflected memory — all version-controllable Markdown.
Memory System
A triple-layer architecture — canonical memories you set, ephemeral moments it captures autonomously, and daily reflections it generates. Together they form a living context that deepens with every session.
--- key: work_context createdDate: 2025-11-01 updatedDate: 2026-06-07 --- Prefers terminal-first workflows. Uses macOS with tmux. Primary languages: TypeScript, Rust. Dislikes verbose explanations. Current project: distributed rate limiter. Deadline: end of sprint.
Set explicitly by you. Loaded into every system prompt at session start. Carries YAML frontmatter with creation and update timestamps.
[auto-captured · expires in 72h] Pattern: user iterates on architecture before writing code Tone: responds well to directness; silence on praise Recurring theme: performance over readability Milestone: shipped without asking for reassurance
Captured autonomously. Records patterns, emotional tone, recurring themes. Influences how madz adapts — without being hardcoded.
--- createdDate: 2026-06-09 updatedDate: 2026-06-09 --- Session synthesis: user pushed hard on performance profiling. Three iterations. Accepted the third without comment. Tone shift detected after first result — less directive, more exploratory. Deadline pressure remains.
Generated nightly by cron at 2 AM. Stored as canonical memories. Auto-installed after onboarding — no configuration required.
Memory Tool Actions
Use Cases
madz is built for environments where you need an agent that can be trusted with real work and held accountable for it.
A self-contained dev environment with a full toolchain, semantic code search, and an orchestrator that can implement, test, and ship features. Spin up one madz per project and you have a dedicated teammate with its own context and memory.
A sandboxed runtime for running untrusted skills and tools, a dedicated security-audit subagent, dependency auditing, and a documented threat model. Run audits in an isolated container without exposing your host.
Medical, financial, and government contexts where auditability, deterministic behavior, and data sovereignty are non-negotiable. Every action is a version-controlled Markdown file, so you can demonstrate exactly what the agent did and why.
A persistent teammate that remembers your context, runs your skills on a schedule, and automates the mundane — with the calm precision of a well-built tool.

Get Started
BSD-3 License · Connect via SSH with account 'madz' · no password · frictionless access
Docker · Quickstart
No password. Madz starts automatically.
docker pull avoidwork/madz:latest docker run -d \ --name madz \ -p 2222:22 \ -v ./memory:/app/memory \ -v madz-checkpoints:/app/memory/checkpoints \ -v ./skills:/app/skills \ -v ./tmp:/app/tmp \ -v ./logs:/home/madz/.cache/madz/logs \ -e OPENAI_API_KEY="your-key" \ avoidwork/madz:latest ssh -p 2222 madz@localhost