Skip to content

MCP server ​

Slide Agent ships an MCP server, slide-agent-mcp, that speaks stdio. It is the integration path for Cursor, Zed, Windsurf, Claude Desktop, and any other MCP client — and unlike a skill, it does not depend on the host implementing a particular skills directory.

The server publishes the composition language and everything around it as resources, so a client that has never heard of Slide Agent can learn to design with it at runtime — start at slide-agent://grammar, which is about 1,300 tokens.

Seven tools by default; --compat-v1 adds the 0.x surface for one more minor release. Both are documented here, 1.x first.


Connect it ​

slide-agent install puts the launcher on your PATH at ~/.local/bin/slide-agent-mcp. Then add it to your client.

Cursor — ~/.cursor/mcp.json

json
{
  "mcpServers": {
    "slide-agent": { "command": "slide-agent-mcp" }
  }
}

Claude Desktop — claude_desktop_config.json

json
{
  "mcpServers": {
    "slide-agent": { "command": "slide-agent-mcp" }
  }
}

Claude Code

bash
claude mcp add slide-agent -- slide-agent-mcp

Zed — settings.json

json
{
  "context_servers": {
    "slide-agent": { "command": { "path": "slide-agent-mcp", "args": [] } }
  }
}

Without installing anything first — any client, at the cost of a download on each launch:

json
{
  "mcpServers": {
    "slide-agent": {
      "command": "npx",
      "args": ["--yes", "-p", "@slide-agent/core@0.9.0", "slide-agent-mcp"]
    }
  }
}

If your client cannot resolve slide-agent-mcp from PATH, use the absolute path — slide-agent doctor prints it, and also tells you whether any host configuration currently references the server.


The loop ​

You direct the design; the engine computes. Six to eight calls, one design review, and no rounds spent hunting overflow.

1. Read the grammar. slides_catalog returns the composition language with worked examples, starter components, recipe one-liners, and presets — under 3,000 tokens, byte-stable for a version, so it caches. slide-agent://grammar is the same page on its own (~1,300 tokens).

2. Write the intent. A brief, a visual concept in plain words, a design language for this deck (palette and roles, type, space, grid, shape, surfaces, texture), components for anything repeated, and a composition per slide in grid units, roles, and named tokens. Recipes are starting points for routine slides; "auto" slides are draft mode and are labelled as such in every verdict.

3. Build. slides_build {deck, intent}. The verdict returns the mechanical state, the adjustments the engine made (contrast repairs, type steps — refuse any of them with a pin), the choices it will not make for you, rhythm notes, and a link to the preview sheet.

Optional and cheap: slides_build {mode: "explore", explore: {designs, slides}} renders up to three design languages across up to four slides side by side, in about a second, before you commit to one.

4. Look. slides_view {deck, what: "sheet"} returns the contact sheet inline. Judge the design against the brief — hierarchy, focus, pacing, craft, anything that reads as generic. Overflow, contrast, and bounds have already been checked; spend the look on design. Record what you decided with slides_build {reviewed: {by, at, notes}}.

5. Edit. slides_edit {deck, ops} with EditOps: change a composition or the design language, answer a choice, shorten to the character budget the engine computed, pin a value so the engine stops adjusting it. Only changed slides are rebuilt.

6. Finalize. slides_finalize {deck, exports} renders with LibreOffice, checks the text survived the render, validates every part against the ECMA-376 schemas, rebuilds the deck from its intent in a clean directory, and exports.

Callers with no model of their own use slides_generate instead: a director model writes the intent, a critic reviews the preview against the brief, and the director revises.

Tools ​

ToolWhat it doesResponse budget
slides_catalogThe composition grammar with worked examples; starter components, recipes, presets, and searches for fonts (fonts: "geometric sans") and icons (icons: "growth")≤ 3,000 tokens
slides_buildValidate, compile the design, solve, fit, check, write deck.pptx, and render a preview sheet. mode: "check" writes nothing; mode: "explore" compares designs≤ 800 tokens
slides_editApply EditOps (intent, design, element pins, package) or, in engine-managed mode, an instruction≤ 500 tokens
slides_viewsheet, slides, crop, rhythm, expand (a recipe as its composition), report, explain≤ 1,500 tokens
slides_finalizeFidelity render, text-survival checks, schema validation, round-trip rebuild, exports≤ 800 tokens
slides_inspectA foreign .pptx or template: layouts, placeholders, brand tokens and locks, per-slide outline≤ 1,500 tokens
slides_generateEngine-managed direction: director, critic, revise. Needs ANTHROPIC_API_KEY and @anthropic-ai/sdk where the server runs≤ 800 tokens

Every tool returns compact JSON. Previews come back as resource links from slides_build, and inline images from slides_view — the call whose purpose is looking. They are previews, drawn from the same measurements the fit engine used, not PowerPoint renders; slides_finalize is what produces those.

What a call costs ​

A verdict is a few hundred tokens whatever the deck's size; the sheet is about 1,850 image tokens for the whole deck, which is what makes looking at twelve slides cheaper than reading two of them. slides_view {what: "report"} pages the full findings when you want them, and explain says why one element looks the way it does — its provenance, fit steps, and every adjustment.

Resources ​

URIWhat it is
slide-agent://grammarThe composition language with worked examples (~1,300 tokens)
slide-agent://catalogGrammar, starter components, recipes, presets as JSON
slide-agent://schema/intentThe full JSON Schema for slide-agent.intent/1 — for validators, not for reading
slide-agent://recipes/{family}/{variant}One recipe: slots, capacity, composition
slide-agent://examples/{name}A complete directed intent with its design language

Paths and roots ​

Every path a tool names is confined to the workspace: the client's MCP roots when it publishes them, SLIDE_AGENT_ROOTS when the operator sets it, otherwise the directory the server was started in. Traversal, absolute escapes, and symlinks out of the workspace are refused, and the error says so. Requests cannot widen policy: remote image fetching stays off unless the operator sets SLIDE_AGENT_ALLOW_REMOTE_IMAGES=1, and build scripts never run from a tool call unless SLIDE_AGENT_ALLOW_SCRIPTS=1.

Compatibility with 0.x ​

slide-agent-mcp --compat-v1 also registers the 0.x tools (slide_agent_run, get_capabilities, get_authoring_contract, plan_presentation, create_presentation, revise_presentation, edit_presentation, render_presentation, validate_presentation, review_presentation, patch_presentation, slide_agent_doctor), their resources, and the author_presentation_scene and revise_presentation_scene prompts. It is there so a host can migrate at its own pace; the 0.x surface is documented below and is frozen.


0.x surface (--compat-v1) ​

The flow that produces good decks ​

It is not prompt → deck. It is build → see → critique → patch, and skipping the seeing is the single most common reason output looks generic.

1. Read what is possible. Call get_capabilities. Its default answer is a summary — what renders here, and whether this installation can source a picture at all. Ask for include: ["canvas"] before you design: that block is the expressive surface, derived from the schemas the engine enforces. Then read the guide sections the deck actually needs with get_authoring_contract; the whole guide is about 8,900 tokens and the router in SKILL.md says which sections matter when.

2. Plan before you place coordinates. Write two visual theses that differ structurally, choose one, and record a sequencePlan — one entry per slide with its narrative job and intended silhouette. If you researched, write the claims and sourceLedger too.

3. Design the deck yourself. Palette, typography, composition, diagrams, and every element's coordinates are your decisions. For anything substantial, write a build script — the same deck as a program runs about a third the length of its NDJSON, and output tokens cost several times what input tokens cost. The full JSON Schemas at slide-agent://contract/schema/<name> are there for validators; the canvas capability block is the cheaper way to learn what you may author.

4. Build it with slide_agent_run, with render on.

5. Look at it with review_presentation. Start with images: "overview": one contact sheet, every slide in order and numbered, which is what makes the deck-level questions answerable — whether the sequence has a shape, whether two slides came out as the same drawing. It costs about one image instead of one per slide. Then open the slides that looked wrong with images: [n], imageDetail: "full". The packet also carries the words read back off the render compared against the deck's own text, the geometry, your declared intent, and questions worth asking.

6. Fix exactly what is wrong with patch_presentation, addressing elements by id. Regenerating the deck to fix a caption discards every decision you are not currently thinking about.

7. Check readiness, not just status, and run roundTrip before you deliver.

jsonc
{
  "request": {
    "command": "create",
    "output": "/absolute/path/deck.pptx",
    "outline": {
      "brief": { "title": "…", "audience": "…", "objective": "…", "…": "…" },
      "narrative": "By the end, the board should approve the migration.",
      "creativeDirection": {
        "name": "Signal through fog",
        "palette": { "background": "0B1020", "ink": "F5F2E9", "accent": "66E3FF" },
        "typography": { "heading": "Georgia", "body": "Aptos" },
        "geometryLanguage": "Hairline routes between few, deliberately placed nodes",
        "visualSystem": {
          "variables": { "signal": "66E3FF" },
          "styles": { "fog-title": { "style": { "fontSize": 48, "color": "F5F2E9", "bold": true } } }
        },
        "avoid": ["rounded corners", "stock photography"]
      },
      "slides": [
        {
          "id": "opening",
          "kind": "statement",
          "title": "One boundary absorbs the complexity",
          "background": "0B1020",
          "canvas": [
            { "id": "title", "type": "text", "x": 0.8, "y": 1.2, "w": 9, "h": 1.6,
              "role": "title", "styleRef": "fog-title",
              "text": "One boundary absorbs the complexity" }
          ]
        }
      ]
    }
  }
}

Then read validation.presentationReadiness in the result before telling the user it worked. packageStatus says the file holds together; readiness says whether the deck is finished, and readinessReasons says why.


Tools ​

ToolRequiredOptional
ToolRequiredOptional
---------
slide_agent_runrequestimages, imageDetail
get_capabilities—include
get_authoring_contract—section, schema
plan_presentationpromptslideCount
create_presentationprompt, outputrender, validate, autoFix, maxRetries, images, imageDetail
revise_presentationinput, output, slide, sceneNdjsonscene, validate, render, images, imageDetail
edit_presentationinput, output, operationsrender, validate, images, imageDetail
render_presentationinput, outputwidth, height, images, imageDetail
validate_presentationinputreport, manifest, previewsDir, render, images, imageDetail
review_presentationinputscene, manifest, slide, from, to, maxSlides, detail, images, imageDetail
patch_presentationinput, output, operationsscene, dryRun, render, roundTrip, validate, images, imageDetail
slide_agent_doctor——

Knowing what is possible before you design ​

get_capabilities and slide-agent://capabilities report what this installation can do. The tool answers with a summary by default and returns any facet in full on request — canvas, images, fonts, rendering, diagrams, charts, layouts, checks, or all. The images block is never summarised away, because it is the one to read before planning a photo-led deck:

json
{ "localPaths": true, "remoteUrls": false, "provider": null,
  "formats": [".png", ".jpg", ".gif", ".webp"] }

remoteUrls: false and provider: null means this installation cannot obtain a picture at all — only embed one already on disk. Design accordingly rather than discovering it after composing eight slides around photography.

provider names a host-installed image resolver: stock search, an internal asset library, an image generator. Slide Agent ships none of these on purpose — see api.md.

Seeing what you built ​

Every tool that can render returns slide previews as image content alongside its JSON result, so a host with no filesystem access of its own can still look at the deck. A model that cannot see its output can only revise from its own assumptions.

What comes back is a choice, and the default is the cheapest correct one:

  • images takes "all", "changed", "none", "overview", or a list of slide numbers. "changed" returns only the slides the command altered — patch knows this from its own diff, revise from its target. When a command cannot tell, it returns everything and says so rather than returning nothing. "overview" composes every slide into one numbered contact sheet.
  • imageDetail is "review" (default, 1024px) or "full" (1568px). The review tier costs roughly half and is sized to judge composition; text fidelity is read from the PDF's text layer, exactly, and never off the image, so the smaller preview costs you nothing there.
  • includeImages still works with its old meaning: true is "all", false is "none".

Up to 20 previews and 12 MB are returned; the text block says how many were withheld, what the call cost, and what the richer option would cost. Where LibreOffice is not installed the previews are schematic SVGs of the deck's geometry rather than rendered slides, and the result says so.

What a call costs ​

Every result carries a tokenBudget:

json
{ "text": 2840, "images": 787, "imageCount": 1, "total": 3627,
  "sessionTotal": 28104, "basis": "estimate" }

The figures are estimates and say so: characters ÷ 4 for text, and (width × height) ÷ 750 after the downscale to 1,568px for images. They exist so an option that saves tokens can be weighed against the one that spends them, which is not a judgement a model can make against an unpublished price list.

slide_agent_run is the one that matters. request.command is create, edit, render, validate, or revise; the rest of the object follows the schema for that command. This is the only tool that accepts a design you authored.

create_presentation takes a prompt and nothing else. It returns a structural draft: your topics, bracketed placeholders where evidence belongs, and no art direction. metadata.provenance reads template-draft. Use it to start a conversation, never to finish one.

plan_presentation returns the same draft as an outline without building a file — useful when you want to show the user a structure before committing.

revise_presentation replaces one slide from the deck's own scene blueprint, leaving every other slide byte-identical. It needs the artifacts/ directory Slide Agent wrote beside the deck. This is almost always better than regenerating.

slide_agent_doctor reports what is installed and what it could verify. Worth calling once if anything behaves unexpectedly.


Resources ​

In three groups: capabilities, the contract descriptor and guide, and one resource per schema.

URITypeContents
slide-agent://capabilitiesJSONThe canvas surface first, then grammars, chart kinds, layouts, checks, fonts, rendering, and how images can reach a slide here
slide-agent://capabilities/canvasJSONEvery element type, property, and treatment the canvas supports, and what stays editable
slide-agent://contractJSONContract version, scene schema id, available schemas
slide-agent://contract/guideMarkdownThe complete authoring guide
slide-agent://contract/guide/<section>MarkdownOne section
slide-agent://contract/schema/<name>JSON SchemaOne schema

Guide sections: role, creative-direction, visual-system, planning, narrative, composition, build-script, canvas, scene, diagrams, data, imagery, accessibility, honesty, review, workflow.

Schemas: outline, brief, slide, canvasElement, creativeDirection, visualSystem, symbol, exploration, sequencePlanItem, claim, hostCapabilities, chart, table, sceneRecord.

Fetch outline when authoring nested slide specs, sceneRecord when authoring the line-oriented NDJSON format, and canvasElement when you only need element geometry and styling.


Prompts ​

author_presentation_scene — argument: brief. Returns the full authoring guide plus the brief, ready to send to a model. Its reply is NDJSON you pass straight to slide_agent_run as sceneNdjson.

revise_presentation_scene — arguments: scene, slide, instruction. Returns replacement records for one slide, preserving the deck's established design system. Pass the reply to revise_presentation.

Use these when your client surfaces MCP prompts as slash commands — they give a user a one-step path to a designed deck.


Paths ​

Every path in a request is a path on the machine running the server. Use absolute paths. There is no upload, no sandbox, and no remote storage: the server writes the deck where you tell it and returns the location.

Remote images are refused unless you set allowRemoteAssets: true on the request, and private and link-local addresses stay blocked even then. Supply local files for images.


Reading results ​

Every tool returns one JSON object as text content.

jsonc
{
  "status": "success",           // or "warning" or "error"
  "primaryOutput": "/…/deck.pptx",
  "deliverables": ["/…/deck.pptx"],
  "artifacts": ["/…/artifacts/…"],
  "slideCount": 8,
  "warnings": [],
  "packageStatus": "pass",
  "presentationReadiness": "review",
  "validation": {
    "packageStatus": "pass",
    "presentationReadiness": "review",
    "readinessReasons": ["Heuristic floor: variety scored 22, below 25. …"],
    "issues": [],
    "heuristics": { "overall": 84, "band": "strong", "dimensions": [/* … */] },
    "fidelity": { "status": "pass", "method": "pdf-text", "confidence": "high", "slides": [/* … */] },
    "artifacts": { "pptx": { "sha256": "…" }, "previews": [/* … */] },
    "suggestedRepairs": [/* what the engine would change, and changed nothing */]
  },
  "errors": [],
  "metadata": { "contractVersion": "0.10", "provenance": "model-authored", "…": "…" }
}

isError is set on the tool result when status is error.

Three things worth checking before reporting success:

  • packageStatus — fail means the file itself does not hold together: a broken relationship, a missing asset, a failed round-trip.
  • presentationReadiness — not-ready means the audience would see a defect. review means something could not be verified. readinessReasons lists what decided it, in order.
  • validation.suggestedRepairs — under the default suggest mode the engine reports what it would change and changes nothing. Read them and decide.

validation.heuristics are engine proxies, not a quality judgement. The report keeps measured facts, heuristics, and reviewer findings apart on purpose.


Common errors ​

CodeMeaning
CONTRACT_VALIDATION_FAILEDYour outline does not match the schema. The message names the exact field path.
REMOTE_ASSETS_DISABLEDAn image URL was used without allowRemoteAssets.
REMOTE_ASSET_BLOCKEDThe URL resolves to a private or link-local address.
SCENE_NOT_FOUNDrevise_presentation or patch_presentation could not find the deck's blueprint. Pass scene explicitly.
PATCH_ELEMENT_NOT_FOUNDA patch named an element the slide does not have. The message lists the ids that exist.
VISUAL_SYSTEM_UNKNOWN_STYLEA styleRef names a style the deck did not declare. The message lists the declared names.
VISUAL_SYSTEM_VARIABLE_TYPEA {"$var":…} landed on a property that cannot accept its type.
RENDER_DEPENDENCY_MISSINGA true render was demanded without LibreOffice and Poppler. By default the server falls back to schematic previews instead.
INPUT_NOT_FOUNDA path does not exist on the server's machine.

Verifying the connection ​

bash
slide-agent doctor

The MCP server check reports whether the launcher exists and whether any host configuration references it. A launcher that works but that nothing calls is the most common misconfiguration, so it is reported separately rather than shown as a pass.

Released under the MIT License.