Skip to content

Repository files navigation

GristCoderMCP — MCP Server for Grist

Turn any Grist document into a full business application using AI.

Grist Coder is an MCP server that connects Claude (or any MCP-compatible LLM) to a Grist document. It ships with a custom widget (IDE + live preview) that lets you build, edit, and deploy HTML/React artefacts — all stored inside your Grist document, no external hosting needed.

MCP Streamable HTTP Python 3.11+ License: MIT Status: Exploratory

Product page · en français — what it does and who it is for, without the implementation details.

MCP server
Name io.github.nic01asFr/gristcoder
Transport streamable-http on /mcp (spec 2025-03-26)
Manifest server.json
Image ghcr.io/nic01asfr/grist-coder
Exposes 35 tools · 8 prompts · 18 resources

Project status: This is an experimental, work-in-progress project developed at Cerema Méditerranée. It works reliably for a single user on localhost, but several features (wizard, sub-agents, chat) are in beta. We publish it to share the approach, gather feedback from the Grist community, and invite contributions toward a complete Grist Coder.

Widget home — artefact selector Map artefact + code editor
Widget home Map + editor

What it does

You say to Claude What happens
"Create a CRM for my contacts" Tables, columns, sample data, a React dashboard, and Grist pages — all wired together
"Add a chart showing sales by month" An HTML artefact with Chart.js, linked to your data via the Grist bridge
"Set up a webhook to notify Slack when a deal closes" A Grist webhook pointing to your Slack endpoint
"Fix the filter on the inventory page" Reads the artefact code, patches it, live-previews the result
"Ship it" Freezes the artefact into the document as a standalone widget — npm imports bundled, no dependency on this server, it keeps working if the server stops and travels with the document

Everything lives inside the Grist document. The artefacts (HTML/React widgets) are stored in an Artefacts table and rendered by the custom widget. Users interact with the finished app — they never see the AI or the code.


Architecture

The 4-layer model

Grist Coder sees a Grist document as a 4-layer application:

Layer          Built with                       What it is
─────────────  ───────────────────────────────  ──────────────────────────────
1. Data        Grist tables + column formulas   The structured information
2. UI          HTML/React artefacts + pages     What users see and interact with
3. Logic       grist_apply, grist_sql, upsert   What the system does on actions
4. Integration grist_webhooks → external APIs   What the document triggers outside

Contextual navigation — progressive tool disclosure

The server does not expose all 35 tools at once. Instead, it tracks a session context that evolves through 6 phases:

qualifying → assessing → designing → building → verifying → done

At each phase, only the relevant tools are visible to the LLM client:

Phase Tools available Purpose
qualifying (15 tools) sessions, wizard, plan, grist read, subagent, chat Understand what the user needs
assessing (19 tools) + canvas_read, screenshot, views_list Audit the existing document
designing (22 tools) + canvas_write, canvas_patch, canvas_type Prototype the UI
building (35 tools) All tools Full construction
verifying (35 tools) All tools Quality check
done (17 tools) = qualifying Delivered, ready for next project

Phase transitions are triggered by plan_update(status=...) which:

  1. Updates the session context
  2. Pushes a notifications/tools/list_changed MCP notification to the client
  3. The client re-fetches the tool list and sees the new set

This prevents the LLM from jumping ahead (e.g., writing code before understanding the schema) and reduces token waste from unused tool descriptions.

Contextual MCP resources

Resources are also organized by phase. Each tool response includes a _next hint and _next_resources pointing to what the LLM should read next:

Phase         Recommended resources
────────────  ──────────────────────────────────────────────────
qualifying    docs/qualification, docs/wizard
assessing     context/{token}, docs/schema, schema-diagram/{token}
designing     docs/schema, docs/artefacts, docs/playbook
building      context/{token}, docs/artefacts, docs/formulas, playbook/{scenario}
verifying     context/{token}, code/{token}

The context/{token} resource is a live snapshot that includes: full schema, FK relationship graph, artefact list, page/section layout, quality audit (_quality: missing visibleCol, empty widgets, orphan tables), and delta between plan and reality (_delta).


The Widget — Grist Custom Widget

The widget (widget.html) is a split-pane IDE served at /:

┌──────────────────────────────────────────────┐
│  Navbar: ‹ ● › Coder          [panel toggle] │
├────────────────────────┬─────────────────────┤
│                        │ Edit bar: [artSelect]│
│   Preview (iframe)     │ Ace editor           │
│   Live render of the   │ Syntax-highlighted   │
│   selected artefact    │ source code          │
│                        │                      │
├────────────────────────┴─────────────────────┤
│  Wizard overlay (when active)                 │
│  ┌─────────────────────────────────────────┐ │
│  │ Multi-card thread: progress, forms,     │ │
│  │ choices, confirmations — all coexist    │ │
│  ├─────────────────────────────────────────┤ │
│  │ Chat bar (appears when LLM sends chat)  │ │
│  └─────────────────────────────────────────┘ │
└──────────────────────────────────────────────┘

Key features:

  • Artefact selector: dropdown to switch between artefacts stored in Grist
  • Live preview: sandboxed iframe with auto-injected Grist bridge (grist.docApi.*, grist.onRecord())
  • Auto-save: edits in Ace are saved back to Grist via browser-side applyUserActions (bypasses server-side WAF)
  • SSE sync: canvas changes from the LLM (via canvas_write/canvas_patch) push to the widget in real-time

Wizard — Interactive AI ↔ User dialogue (beta)

The wizard is an overlay system that lets the LLM interact with the user directly through the widget, without requiring the user to type in the LLM chat interface.

How it works

  1. The LLM calls canvas_wizard(step) with a step definition
  2. The server pushes the step via SSE to the widget
  3. The widget renders an interactive card (form, choice, confirmation...)
  4. The user responds in the widget
  5. The response is sent back to the server via POST /wizard/{token}
  6. The canvas_wizard tool call unblocks and returns the user's answer to the LLM

Card types

Type Behavior Use case
input Blocking Free-text collection with suggestion chips
choice Blocking Category selection (clickable cards)
form Blocking Structured data collection (text, number, select, toggle)
confirm Blocking Markdown preview + accept/reject
progress Non-blocking Build progress with status indicators
info Non-blocking Contextual information
preview Blocking Live iframe preview + interaction
data-import Blocking Fetch external API → preview table → import to Grist

Multi-card thread

Multiple cards coexist in the overlay — a non-blocking progress card can stay visible while a blocking form card collects input. Cards are managed independently:

  • canvas_wizard(id="build-progress", type="progress", ...) — persistent progress tracker
  • canvas_wizard(id="confirm-plan", type="confirm", ...) — blocking validation
  • canvas_wizard_close(card_id="confirm-plan") — close one card
  • canvas_wizard_close() — close everything

Current limitations (beta)

  • The wizard UX is functional but rough — styling and transitions need polish
  • Card layout on mobile/small viewports is not optimized
  • The data-import card type works but error handling is minimal
  • Chat integration (chat_reply / wait_for_chat) is basic — no message history persistence
  • Sub-agent delegation (subagent_call) depends on MCP client sampling support; fallback mode works but is less reliable

Quick start

Option A — Python (development)

git clone https://gitlab.cerema.fr/mcp/gristcoder_mcp.git
cd grist-coder-mcp

python -m venv .venv
# Windows
.venv\Scripts\activate
# macOS/Linux
source .venv/bin/activate

pip install -r requirements.txt
cp .env.example .env        # edit HOST_URL if needed
uvicorn grist_coder:app --port 8742 --reload

Option B — Docker

git clone https://gitlab.cerema.fr/mcp/gristcoder_mcp.git
cd grist-coder-mcp

cp .env.example .env
docker compose up -d

Verify

curl http://localhost:8742/health
# → {"ok": true, "version": "5.14", "tools": 35, ...}

Setup

1. Add the widget to Grist

In your Grist document:

  1. Add a new Custom Widget
  2. Set the URL to http://localhost:8742/
  3. Grant Full document access when prompted

The widget registers itself automatically — no API key needed on the widget side.

2. Connect Claude Desktop

Edit %APPDATA%\Claude\claude_desktop_config.json (Windows) or ~/Library/Application Support/Claude/claude_desktop_config.json (macOS):

{
  "mcpServers": {
    "grist-coder": {
      "type": "http",
      "url": "http://localhost:8742/mcp",
      "headers": {
        "Authorization": "Bearer YOUR_GRIST_API_KEY"
      }
    }
  }
}

Get your Grist API key from Grist → Profile → API.

Restart Claude Desktop after editing.

3. Connect Claude Code (CLI)

Copy .mcp.json.example to .mcp.json and fill in your Grist API key:

cp .mcp.json.example .mcp.json
# edit .mcp.json with your key

4. Run it hosted (and what guards it)

The server is not localhost-only. charts/grist-coder/ deploys one pod per user on SSPCloud Onyxia. Online, it exposes the same MCP surface, with guards that do not depend on the client behaving:

Guard Env What it does
Pod gate APP_AUTH_TOKEN Bearer token required on /mcp, /register and /llm-proxy, compared with hmac.compare_digest. Empty (local default) = no gate
Owner lock (TOFU) OWNER_LOCK The first real Grist account to register pins the pod; any other uid gets a 403. Unresolved fallback identities are never pinned — they would lock the pod to nobody
OAuth 2.1 connector PUBLIC_URL The pod becomes its own authorization server. The Grist key is given once at consent and stays server-side; the client only ever holds an opaque gco- token that expires and can be revoked
Error scrubbing _scrub_secrets() masks auth= and Bearer in every message returned to a client. The key used to leak through the httpx request URL
SSRF closed by default LLM_PROXY_ALLOWED_HOSTS /llm-proxy only reaches allowlisted hosts. Empty list disables the endpoint entirely
Key containment /llm-config returns base, model and cle_serveur: boolnever the key. The in-widget agent never sees it
Session hygiene SESSION_TTL Idle sessions purged (24 h default); SSE fan-out routed per uid, so no cross-user delivery

The owner lock matters most on a public document: without it, any visitor authenticating anonymously would land in the same uid bucket and see each other's sessions.


MCP Tools (35)

Sessions

Tool Description
sessions_list List open Grist documents — call first
session_open Open a document without a browser, from the caller's Grist key
session_select Pin a default session for this connection
session_info Full context: doc, tables, artefacts, canvas state

Every document-scoped tool also accepts an optional token argument. Routing is carried per call, not by shared server state — so several agents or browser tabs can work on different documents in parallel.

Plan & context

Tool Description
plan_update Update the work plan, transition phase, trigger tool list change
savoir_faire Proven recipes and traps from the Grist ecosystem — returns 2 bounded units, not a chapter. Call before the first artefact of a build, or when an error names something unfamiliar

Canvas (artefact editor)

Tool Description
canvas_select Switch to a different artefact (read-only)
canvas_read Read the current artefact source code
canvas_write Write/replace the full artefact code
canvas_patch Surgical find-and-replace (old_strnew_str)
canvas_exec Execute Python canvas (sandboxed subprocess, 10s timeout)
canvas_screenshot Capture the rendered artefact as PNG
canvas_type Change artefact type (html, react, markdown, mermaid, python, sql, svg...)

Interactive — wizard / chat (beta)

Tool Description
canvas_wizard Show an interactive card in the widget (choice, form, confirm, progress...)
canvas_wizard_close Close a specific card or the entire overlay
canvas_context_update Non-blocking context card in the overlay
chat_reply Send a chat message — optionally wait for user response
wait_for_chat Wait for user input without sending
subagent_call Delegate to a specialized sub-agent (data-architect, ui-designer...)

Artefact management

Tool Description
artefact_init Create the Artefacts table if missing (idempotent)
artefact_publish Freeze a finished HTML artefact as a fully standalone widget stored inside the doc (gallery "Custom widget builder") — no MCP server needed at runtime. Bundles npm imports automatically (esbuild), and refuses any artefact that would depend on the pod. mode="app" publishes a shell that lazily mounts the app/… rows as screens of a multi-screen application

Grist — Read

Tool Description
grist_schema Document schema (tables, columns, types)
grist_records Read records from a table (with optional filter/limit)
grist_sql Run arbitrary SQL on the document

Grist — Write

Tool Description
grist_records_add Insert new records
grist_records_patch Update existing records by ID
grist_upsert Upsert on a business key (require + fields)
grist_apply Low-level Grist UserActions (AddTable, AddColumn, etc.)
grist_validate Pre-flight — validates a list of UserActions without writing. Catches the classic Ref: to a table created later in the same batch, invalid action shapes, missing visibleCol. Call before any multi-AddTable grist_apply

Document / Pages

Tool Description
grist_views_list List pages with their widget sections
grist_view_create Create a page with a grid + optional artefact widget
grist_view_add_widget Add a widget section to an existing page
grist_section_configure Configure a custom widget section (artefact, linking...)
grist_doc_create Create a new empty document in a workspace, add the Coder widget page, return the clickable link. Uses the caller's Grist key — no server configuration required

Webhooks

Tool Description
grist_webhooks CRUD webhooks (list, create, update, delete)

HTTP Endpoints

Method Route Purpose
POST /mcp MCP JSON-RPC (tools, resources, prompts)
GET /mcp SSE stream — live canvas updates to widget
DELETE /mcp Close an SSE session
POST /register Widget auto-registration
POST /wizard/{token} Wizard form response from widget
POST /chat/{token} Chat reply from widget, unblocks wait_for_chat
POST /apply-result/{token} Result of a browser-side applyUserActions round-trip (WAF bypass)
POST /art-diag/{token} Render diagnostics from the artefact iframe (exceptions, failed resources)
GET /llm-config What the pod knows about the LLM (base, model, whether it holds a key — never the key itself)
POST /llm-proxy/{path} CORS-free relay to an allowlisted LLM host; injects the pod key when the browser sends none
GET /harness/{file} Browser-side agent modules
POST /webhook-receive/{docId} Grist webhook receiver → SSE fan-out (production, public URL required)
GET / Serves the custom widget
GET /survey/{docId}/schema-check Survey Manifest conformity check for a document
GET /health Health check — source of truth for version and counts

OAuth 2.1 connector (hosted deployments only, active when PUBLIC_URL is set)

Method Route Purpose
GET /.well-known/oauth-authorization-server Authorization server metadata
GET /.well-known/oauth-protected-resource Protected resource metadata
POST /oauth/register Dynamic client registration (permissive — a strict one breaks "cannot register" in clients)
GET/POST /oauth/authorize Consent screen: the user pastes their Grist key once
POST /oauth/token Exchanges the code for an opaque gco- token (PKCE S256)

MCP Resources

Resources provide contextual documentation and live data to the LLM.

Static documentation

URI Content
grist-coder://docs/schema Column types, formulas, visibleCol recipes
grist-coder://docs/artefacts Artefact templates, Grist bridge API, DSFR components
grist-coder://docs/playbook Page creation sequences, linked sections, decision tree
grist-coder://docs/app-patterns Multi-widget patterns: navigation, sync, routing
grist-coder://docs/formulas Grist Python column formulas (isFormula, lookupOne...)
grist-coder://docs/wizard Wizard card schema and types
grist-coder://docs/qualification App categories, architecture templates, completeness criteria
grist-coder://docs/services-geo Geocoding, maps (BAN, OSM, IGN, Leaflet)
grist-coder://docs/services-data Open data (SIRENE, DVF, data.gouv, API Geo)
grist-coder://docs/services-ai AI patterns (sync, async webhook, bridge, sub-agent)

Live document context (templates)

URI Content
grist-coder://plan/{token} Persistent work plan + phase + _next_step hint
grist-coder://context/{token} Full document snapshot: schema, FK graph, artefacts, pages, _quality audit, _delta plan vs reality
grist-coder://schema-diagram/{token} Auto-generated Mermaid ER diagram
grist-coder://code/{token} Source code of all artefacts
grist-coder://context/{token}/page/{page_id} Page-level context: source table, linked sections, configured artefact
grist-coder://playbook/{scenario} Scenario guide: dashboard, fiche, table, full-app, master-detail
grist-coder://examples/{domain} Schema + sample data for: crm, rh, stock, projets, immobilier, association, restaurant, formation

How artefacts work

Artefacts are stored in a Grist table called Artefacts:

Column Purpose
Nom Unique name (used as key)
Type html, react, app, markdown, mermaid, python, sql, svg
Code The source code
Description What this artefact does

The bridge contract

Every artefact is rendered in a sandboxed iframe with a Grist bridge injected into it. This is the contract an artefact can rely on — and the reference for it, since the older documents in this ecosystem still describe the pre-helper API.

Reading. grist.docApi.loadTable(t) returns an array of row objects, [{id, Col, …}], ready for rows.map(…). It does not return a table-keyed envelope: result.MaTable is undefined, and an || [] after it turns that mistake into a silently empty screen. The array now answers to .MaTable by returning itself and logging the fault, so the screen survives and the diagnostic names what to fix — but write rows directly. grist.docApi.fetchTable(t) is still there and still columnar ({id: [...], Col: [...]}), for when that shape is what you want.

Writing. addRow(t, {Col: v}), updateRow(t, id, {Col: v}), deleteRow(t, id). Each builds the correct UserAction and returns the refreshed rows, so a write is followed by a render without a second round trip. applyAndFetch(actions, table) does the same for hand-written actions. In a browser session use BulkAddRecord, never BulkAddOrReplaceRecord.

Helpersgrist.util.*, so nobody re-derives them: toRows(d) (columnar → objects), toDate(ts) / fromDate(d) (Grist stores dates in seconds; printing the raw value gives 1704067200 instead of a date), refIds(l) / toRefList(a) (a RefList is prefixed 'L'), esc(t) (use it before putting a cell value in innerHTML).

Who is lookinggrist.user is {id, email, nom} for the connected Grist user. Use it instead of hardcoding a row id for "the current user". It is for display: real enforcement is Grist access rules, evaluated server-side where user.Email and user.Access are available, and the widget only ever receives rows the user may see.

grist.onRecord() / onRecords() remain available for artefacts driven by the widget's selected table.

Artefact types

Type Rendered as Use case
html Raw HTML in iframe Dashboards, forms, custom UIs
react React 18 + Babel (CDN) Complex interactive components
app React + multi-view router Full single-page applications
markdown Rendered Markdown Documentation, reports
mermaid Mermaid diagrams Flowcharts, ER diagrams
python Executed via canvas_exec Data processing scripts
sql Executed via grist_sql Analytical queries
svg Inline SVG Icons, illustrations

Render diagnostics — the correction loop

An artefact that fails renders a blank page, and the agent that just wrote it learns nothing. Without feedback it either declares success on broken code, or needs a human to open the console.

A probe is injected into every rendered artefact — before the Grist bridge and before the artefact's own code, so initialisation errors are caught too. It reports back exceptions (message, line, column, first stack frames), unhandled promise rejections, console.error, resources that failed to load, and the state of the render: element count, text length, canvas/svg presence.

That last one matters most, and it is not obvious. An error is not required for an artefact to be broken: a screen whose script fails early shows its markup and nothing else, without throwing anything. A headless render would report "page loaded, zero errors". Measuring what was rendered, not just what crashed, catches it.

canvas_write waits briefly for the result and returns it in the same response:

{
  "ok": true, "sha": "4d9f44f7",
  "diagnostic": {
    "erreurs": [{"message": "Uncaught ReferenceError: calculerTotal is not defined",
                 "ligne": 8, "colonne": 13, "pile": "..."}],
    "resume": "Artefact en echec au rendu : 1 exception(s)."
  },
  "_next": "Corriger puis reecrire — le diagnostic revient a chaque canvas_write."
}

The agent fixes and rewrites; the diagnostic disappears. No screenshot, no console, no human in the loop.

Two limits. It needs a widget open on that document — nothing renders without a browser, so there is nothing to observe. And it covers the initial render: an error triggered by a later click is captured but no longer awaited — read it back with session_info, which returns the last diagnostic.


Publishing — standalone widgets

artefact_publish freezes an artefact into the document itself, in the section options of a gallery widget. The published widget does not depend on this server at runtime: it survives the pod being stopped and travels with the document (copies, exports).

That promise is verified, not merely stated. Publication is refused when the code references the pod — /ai-proxy, /llm-proxy, /webhook-receive, or the pod's own URL — because such a widget would die with the server. Use grist_view_create instead for artefacts meant to stay served by the pod.

npm imports are bundled automatically. An artefact doing import { render } from 'preact' cannot run in a browser as-is — it would render blank. On publish, the server resolves the imports with esbuild (a static Go binary in the image, no Node) and inlines everything. Packages are fetched from the npm registry on demand and cached. Transitive dependencies are resolved by retrying on esbuild's own "Could not resolve" errors, which converges without reimplementing npm.

JSX works inside a html artefact — the type check applies to Artefacts.Type, not to the content.

External CDNs still work in a published widget (measured: a <script src> pointing at jsDelivr does execute). They are reported, not blocked: the widget then depends on that CDN at runtime and breaks if it becomes unreachable. Bundling is therefore a robustness choice, not a technical necessity.

mode="app" — a multi-screen application in one widget

Rows named app/Accueil, app/Detail… become screens. Publishing with mode="app" installs a small shell that lists them, mounts one on demand, and gives each screen app.navigate(), app.emit() / app.on(), app.setState() plus the Grist API relayed over postMessage.

Two properties make this worth it. Section metadata stays tiny — Grist downloads all _grist_* tables in full on every document open, so a monolithic artefact costs that weight to everyone, every time. And screens are fetched one SQL query at a time: a screen nobody visits is never downloaded.

Measured on a three-screen application against a monolithic artefact of comparable content: 6.5 KB of section metadata instead of 755 KB, 237 bytes transferred at open, and an unvisited screen never fetched.

Editing a screen needs no republication — change the row, reload the page.


Browser-side agent (harness)

Everything above assumes an external MCP client — Claude Desktop, Claude Code — driving the tools. The harness is the other way round: an LLM agent that runs inside the widget, in the browser, and calls the same MCP tools over the same /mcp endpoint. The document then builds itself from Grist alone, with no desktop client in the loop.

Seven vanilla-JS modules under harness/, served by the pod at /harness/{file}:

Module Role
boot.js Load order and wiring into the widget
config-panel.js LLM configuration + the robot button in the navbar
llm-client.js Transport to /llm-proxy, timeouts, scrubbed errors
agent-loop.js The turn loop: prompt → tool calls → observations
mcp-tools.js Tool discovery and invocation against /mcp
agent-memory.js Conversation memory across turns
render-bridge.js Who drives the render pane — local (agent) or sse (server)

Configuration comes from the pod

The LLM key used to live in the browser's localStorage, retyped on every machine. It doesn't have to: the pod usually holds one already, and knows which base and model to use. GET /llm-config reports that — never the key itself, only whether one exists. The panel prefills base and model, and stops demanding a key when the pod can supply it. When it can't, it says so plainly instead of refusing generically.

LLM_BASE_URL=https://llm.lab.sspcloud.fr/api
LLM_MODEL=qwen3-6-35b-moe
LLM_PROXY_ALLOWED_HOSTS=llm.lab.sspcloud.fr,albert.api.etalab.gouv.fr
LLM_API_KEY=...                 # or LLM_AUTO_FROM_DATALAB=true on SSPCloud

Reusing the Onyxia profile key

Onyxia asks for the LLM key once, in the user profile (AI Assistant tab), and resolves {{userProfileValues.aiAssistant.apiKey}} at launch time, from its own UI. The chart declares that placeholder, so launching this service from the Onyxia catalogue fills the key field by itself.

Installing with helm install gets none of that — nobody resolved the placeholder. This is worth stating because the failure is silent: the pod simply has no key. The chart alone is not enough, and a chart carrying the placeholder proves nothing about a CLI-installed release.

So LLM_AUTO_FROM_DATALAB=true recovers the key where it actually sits. A service launched from the Onyxia UI ends up with the key written literally into its manifest (OPENAI_API_KEY), and the pod reads the workloads of its own namespace — the user's own space, with the edit ClusterRole the chart already grants. Guards, each earned:

  • our own workload is skipped — it holds the emptiness we are trying to fill;
  • only literal values count, never a valueFrom pointing at a Secret we may not read;
  • when the source declares its own LLM base, it must match ours — otherwise another service's OpenAI key gets offered to SSPCloud, which answers 401;
  • candidates are validated against {base}/v1/models before being adopted, newest workload first. A key frozen in a manifest ages: the one sitting in a four-day-old StatefulSet was already answering "session has expired". Announcing cle_serveur: true for a dead key is worse than admitting there is none — the widget stops asking for a key and the failure surfaces later, somewhere unrelated.

The log distinguishes nothing found from found but refused, and names the source: the remedy differs. Relaunching any service from the Onyxia UI refreshes the key frozen in its manifest; otherwise set LLM_API_KEY on the pod.

These keys age by design. The region declares its AI gateway as OIDC-backed (oauthProvider: oidc, a token-exchange bridge) and states that credentials are injected at each service start. A rejected key therefore means "stale", not "wrong". So a key the pod adopted at boot can die mid-life, and a cache would keep serving it until the next restart — on a 401/403 from the gateway the pod forgets what it thought it knew and searches again on the following call.

Only the key is taken from the profile. The profile also shows a Default model and an API base URL, but the profile's default model is not guaranteed to exist on the gateway the key opens: measured here, the profile said devstral-2:123b while the gateway answered Model not found and served qwen3-* / gemma4-* instead. Wiring the model through would reintroduce exactly the opaque failure this section exists to prevent. The corresponding Onyxia placeholder paths are also unverified — a placeholder pointing at a path that does not exist resolves to empty, which would silently wipe a working base URL at launch.

Some Onyxia versions instead materialise the profile as a *secretassistant Secret; that path is still tried, second, and validated the same way.

The Arreter button in the panel stops a running agent and hands the render pane back to the server driver.

Choosing the model — check the catalogue first

The SSPCloud catalogue changes under you. gemma3-27b-it, the previous default here, no longer exists, and the only symptom is a flat {"detail":"Model not found"} from the proxy — nothing points at the model name. List what is actually served before configuring anything:

curl -s "$LLM_BASE_URL/v1/models" -H "Authorization: Bearer $LLM_API_KEY" | jq '.data[].id'

Measured on llm.lab.sspcloud.fr (emits a structured tool_calls, and uses the tool result on the next turn instead of calling it again):

Model Tool-calling
qwen3-6-35b-moe works — current default
qwen3-cursor works — leans towards code
gemma4-26b-moe works
qwen3-vl unusable: server started without --enable-auto-tool-choice

Honest status

Native tool-calling is verified end to end against this service: the model emits the call, the loop feeds the result back, and the model answers from it. What has not been exercised is a long build — many turns, many tools, an actual application produced from a blank document. Treat the harness as working-but-young: the mechanism holds, its stamina is unmeasured.


Authentication model

Widget (browser)                    Server
──────────────────                  ──────
grist.docApi.getAccessToken()  →  POST /register  →  uid:{userId}
                                                   →  session token gc-xxxxxx

Claude Desktop                      Server
──────────────────                  ──────
Authorization: Bearer <api_key>  →  GET /api/profile/user  →  uid:{userId}

Both paths resolve to the same uid:{userId} identity. Sessions are shared between the widget and Claude — they see the same artefacts and canvas state.


Security

  • canvas_exec runs arbitrary Python in a subprocess — only expose to trusted users
  • Do not expose this service publicly without additional authentication
  • WEBHOOK_SECRET (optional, recommended in production): set in .env and configure the same secret in Grist webhook headers
  • API keys are never returned in tool responses

What works, what doesn't (honest status)

Stable (single user, localhost)

  • MCP server core: tool dispatch, resource serving, SSE streaming
  • Grist CRUD tools: schema, records, sql, apply, upsert, webhooks
  • Canvas tools: read, write, patch, screenshot, type detection
  • Widget: artefact editing, live preview, auto-save, Grist bridge injection
  • Authentication: widget auto-registration + Claude Desktop API key
  • Docker deployment

The server architecture supports multiple users: per-user sessions (uid:{grist_user_id}), isolated SSE streams (events filtered server-side), and per-user Grist API credentials. However, it has only been tested in single-user local deployments. Multi-user and remote deployments are untested and would require additional hardening (HTTPS reverse proxy, rate limiting, canvas_exec sandboxing).

Errors that used to be silent — now named

Three failures cost the most time here, and all three were silent in the same way: the thing that broke was never the thing the message described. They are worth stating, because the fixes are the reason the current version behaves.

A Grist 500 with no explanation. httpx keeps only Server error '500 …' for url and discards the body — which is exactly where Grist says what is wrong. Every failure looked identical, so an agent told to retry a transient error retried a permanent one: one run burned its entire 32-iteration budget that way. The body now reaches the caller, and an error naming its own cause is no longer replayed.

A document made unopenable by a Choice column. An agent passed widgetOptions as an object — the natural form over JSON-RPC. It crossed Grist's Python sandbox, came back as {'choices': [...]} in repr() form, and the frontend could no longer parse it: the document stopped opening entirely. Meta fields that Grist stores as JSON strings are now serialised on the way in, and a string that opens like JSON without being JSON is refused before anything is written.

An expired widget token shadowing a valid API key. The query token was sent whenever present, so once the document became unopenable — no browser, no fresh token — every tool answered 401, including the ones needed to repair it. A session holding an API key now uses only that.

Beta — functional but needs work

  • Render diagnostics: exceptions, rejections and failed resources come back reliably. The "rendered nothing without throwing" heuristic is cruder — it flags a body with almost no elements and no text, which can produce a false positive on a deliberately minimal artefact.
  • Wizard system: multi-card overlay works, but UI polish is lacking. Transitions between phases can feel abrupt. The data-import card type is useful but fragile with malformed API responses.
  • Chat integration: chat_reply and wait_for_chat work. History is kept server-side (a 200-message ring per session) and replayed on SSE reconnect, so refreshing the widget no longer loses it — but a pod restart does, sessions being in memory.
  • Sub-agents: subagent_call works when the MCP client supports sampling/createMessage (Claude Desktop). Fallback mode (for Claude Code and other clients) works but the LLM must manually adopt the sub-agent role, which is less reliable.
  • Contextual tool filtering: the phase-based tool disclosure works correctly, but the phase transitions could be smoother — sometimes the LLM needs a tool that's not yet available in the current phase.

Experimental / incomplete

  • Browser-side agent (harness): loads, configures itself from the pod, calls tools and consumes their results, and stops cleanly. What is untested is stamina — a full multi-turn build from a blank document. See the harness section above.
  • Plan persistence: plans live in-memory (session), not in Grist. Server restart = plan lost. We intend to store plans in a Grist table.
  • Webhook receiver (/webhook-receive/{docId}): works in production with a public URL, but not usable on localhost without a tunnel (ngrok, etc.)
  • Styling is a floor, not a theme: every artefact receives a small stylesheet written entirely in :where() — font, margins, a legible table, usable buttons and inputs. Zero specificity, so any rule the artefact writes wins. It exists because an agent given no instruction about appearance produces raw HTML, and a demo in raw HTML reads as broken. DSFR is not injected — artefacts that want it declare the two jsDelivr <link> tags themselves.

Known limitations

  • The documentation this server carries does not reach the agent it embeds. Eleven docs/* resources and eight example domains are served as MCP resources — reachable by Claude Desktop or Claude Code, invisible to the browser-side harness, which speaks only tools/list and tools/call. It reads exactly one resource, context/{token}, to drive the plan banner. Measured consequence on a real build: the agent rewrote its own toDate although grist.util.toDate is injected into every artefact, used grist.util nowhere at all, produced no styling, and hardcoded the current user as row 1. It was not disobeying — it had the system prompt, the tool descriptions and the document schema, nothing else.
  • A second corpus is not here at all. Widgets Grist/skills/ holds ~2200 lines of proven widget patterns — base template, modals, toasts, filters, CSS theme variables, status mappings — referenced only from a code comment.
  • No access control. Nothing in this server knows Grist access rules exist: no tool, no resource, no prompt line. An app built here can present per-role screens while enforcing nothing, and an artefact has no way to learn who is looking at it — hence the hardcoded row 1 above. Both halves need work: grist.user in the bridge (display), and access rules on the data (enforcement).
  • Single file architecture (grist_coder.py, ~8200 lines) — intentional for deployment simplicity, but makes contribution harder
  • In-memory sessions — no horizontal scaling, no persistence across restarts
  • The widget is vanilla JS (~2000 lines) — no framework, no build step, which keeps it simple but limits maintainability
  • canvas_exec and /run execute arbitrary Python — this is a feature for trusted environments, a risk for public ones
  • A published widget still depends on the gallery builder (@berhalak/custom-widget-builder, served from GitHub Pages). "No dependency on the MCP server" is exact; "no dependency at all" is not. Bundling removes the library CDN, not that one.
  • Instances behind a WAF may refuse JSX source on server-side writes. On grist.numerique.gouv.fr, a payload containing bare HTML tags (the JSX signature) is answered 403 — escaping </> does not help, the WAF normalises JSON escapes. Writes fall back to the browser automatically when a widget is open; otherwise write the source with h(...) / React.createElement, which bundles just as well.
  • canvas_screenshot needs an open widget by nature — the iframe captures itself. It is the only tool session_open cannot serve.

Project structure

gristcoder_mcp/
├── grist_coder.py       # MCP server (single file, ~8200 lines)
├── widget.html           # Grist custom widget (IDE + preview + wizard)
├── harness/              # Browser-side LLM agent (see below) — 7 modules
│   ├── boot.js           #   entry point, called once the widget is registered
│   ├── agent-loop.js     #   the loop: tool definitions, execution, specialists
│   ├── llm-client.js     #   LLM calls (through /llm-proxy)
│   ├── mcp-tools.js      #   exposes the MCP tools to the agent
│   ├── agent-memory.js   #   conversation memory
│   ├── render-bridge.js  #   rendering into the widget
│   └── config-panel.js   #   configuration UI + "Lancer" button
├── requirements.txt      # Python dependencies
├── Dockerfile            # Container image
├── docker-compose.yml    # One-command deployment
├── .env.example          # Environment template
├── .mcp.json.example     # Claude Code MCP config template
├── ARCHITECTURE.md       # How it works — capabilities, widget types, user guide
├── CONTRIBUTING.md       # Developer guide — code map, how to add tools/resources/prompts
├── CLAUDE.md             # AI assistant instructions (for contributors using Claude Code)
├── LICENSE               # MIT
└── README.md             # This file

Compatible Grist instances

Tested with:

Any Grist instance exposing the standard REST API should work.


Tech stack

  • Python 3.11+ · FastAPI · uvicorn · httpx
  • MCP transport: Streamable HTTP (spec 2025-03-26)
  • Widget: Vanilla JS + Ace Editor + Grist Plugin API
  • No database: sessions are in-memory, all persistent state lives in Grist

Contributing

Where this project lives

Development happens on GitLab CEREMA — that is where branches, merge requests and CI run. The GitHub repository is a mirror, kept for visibility and for anyone outside the CEREMA network. It is not the working copy: an issue or pull request opened there may go unnoticed, and a commit pushed there would be overwritten by the next mirror sync.

If you cannot reach GitLab CEREMA, open the discussion on the GitHub mirror anyway and say so — we will carry it across.

This project is exploratory and we welcome contributions — whether it's bug reports, feature ideas, or pull requests. We're particularly interested in:

  • Wizard UX improvements — better card styling, animations, mobile support
  • Session persistence — storing plans/state in Grist tables instead of memory
  • Styling — a minimal zero-specificity floor is injected into every artefact so that an unstyled one is still legible; making it themeable (and optionally DSFR) is open work
  • Testing — there are currently no automated tests
  • Documentation — usage guides, video demos, example workflows

How to contribute

  1. Fork the repo on GitLab CEREMA
  2. Create a feature branch (git checkout -b feat/my-feature)
  3. Make your changes in grist_coder.py and/or widget.html
  4. Test with a real Grist document
  5. Submit a merge request

Code conventions

  • grist_coder.py is a single file by design — do not split it
  • New tools: add to TOOLS[] list + handle in call_tool()
  • New prompts: add to PROMPTS[] + handle in _prompt_messages()
  • New resources: add to STATIC_RESOURCES[] + handle in _read_resource()

License

MIT — Nicolas LAVAL, Cerema Méditerranée

About

Grist Coder is an MCP server that connects Claude (or any MCP-compatible LLM) to a Grist document. It ships with a custom widget (IDE + live preview) that lets you build, edit, and deploy HTML/React artefacts — all stored inside your Grist document, no external hosting needed.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages