Agent contract · v1

Teach your agent
to roll local.

A precise operating manual for LLMs, coding agents, and automation that need private models, tool calls, images, video, or speech through Tapioca.

DISCOVER, DON'T GUESS●LOOPBACK BY DEFAULT●TOOLS STAY PERMISSIONED●OUTPUTS STAY LOCAL

01 · Deterministic quickstart

Six moves from detection to cleanup.

The catalog—not memory—is the source of truth. Select the narrowest execution surface and leave the machine as you found it.

  1. 01

    Detect

    tapioca version

    Stop and install Tapioca if this command is unavailable.

  2. 02

    Inspect

    tapioca list && tapioca catalog update && tapioca catalog

    Refresh verified recipes, then prefer an installed, task-compatible model.

  3. 03

    Acquire

    tapioca pull MODEL

    Tell the user before a large model download. Never accept gated terms for them.

  4. 04

    Serve

    tapioca serve MODEL --host 127.0.0.1 --port 11435

    Keep the API on loopback and wait for /health.

  5. 05

    Act

    POST /v1/responses

    Validate model tool calls through the host agent's permissions.

  6. 06

    Clean up

    stop the server you started

    Leave unrelated local models and files untouched.

02 · Compatibility APIs

One local server.
Three familiar protocols.

Start on http://127.0.0.1:11435, wait for health, then use the request shape your client already understands.

GET/health

Wait for the active model to finish loading.

GET/v1/models

Confirm the model exposed by this server.

POST/v1/chat/completions

OpenAI chat, streaming, and tool calls.

POST/v1/responses

OpenAI Responses input and function tools.

POST/v1/messages

Anthropic Messages and Claude Code compatibility.

curl http://127.0.0.1:11435/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer tapioca-local" \
  -d '{
    "model": "glm-4.7-flash:q8_0",
    "messages": [
      {"role": "user", "content": "Say hello in one sentence."}
    ],
    "stream": false
  }'

03 · Tool calling

Tapioca transports.
Your agent decides.

Models propose tool calls; Tapioca never executes them. Validate function names and arguments, apply the host's approval policy, return a matching tool result, and cap the loop.

  1. 1Send JSON function definitions
  2. 2Validate the returned call
  3. 3Execute only an approved host tool
  4. 4Return the result with its call ID
{
  "type": "function",
  "function": {
    "name": "read_project_file",
    "description": "Read a UTF-8 file in the workspace",
    "parameters": {
      "type": "object",
      "properties": {"path": {"type": "string"}},
      "required": ["path"],
      "additionalProperties": false
    }
  }
}

04 · Bundle-aware media

Ask for a model.
Not a pile of weights.

Agents treat catalog IDs as bundle contracts. MiniMax-H3 resolves the correct MPS or CUDA bundle, while Tapioca privately owns the engine graph, adapter ordering, and output path.

MiniMax-H3 is subject to its community license, which excludes the US, EU, UK, and Republic of Korea absent separate authorization. An ungated download or local acknowledgement is not permission to use it there.

LoRA automation uses the provider-neutral library: list installed references first, inspect or pull from Hugging Face, Civitai, or ModelScope, and import existing files into a managed local reference.

For computer-to-computer migration, preserve complete adapter snapshots and rebuild copied base-model registration with tapioca pull MODEL. The import and transfer guidegives agents and humans the exact safe workflow.

1tapioca catalog update && tapioca catalog

Refresh verified recipes, then confirm platform, memory, and the exact model ID.

2adapter list

Reuse its canonical reference when the required LoRA is already installed.

3inspect / pull / import

Acquire from Hugging Face, Civitai, ModelScope, or a verified local file.

4ordered --adapter

Apply transformer LoRAs in command order; never infer compatibility by extension.

5verify output

Wait for exit, preserve the returned path, and inspect video plus audio streams.

!gated model

Stop for user review. Never accept provider terms or expose an access token on the user's behalf.

tapioca pull minimax-h3
tapioca adapter list
tapioca adapter inspect hf://OWNER/REPOSITORY
tapioca adapter inspect civitai://MODEL_ID/VERSION_ID
tapioca adapter import ./adapter.safetensors --base minimax-h3 --name local-adapter
tapioca video minimax-h3 \
  --prompt "A cinematic tracking shot" \
  --adapter 'local://local-adapter#adapter.safetensors@0.8' \
  --preset low-memory \
  --output adapted.mp4

05 · Coding clients

Launch the agent.
Keep its real profile clean.

Tapioca creates isolated configuration under TAPIOCA_HOME/launch. The client must already be installed on the machine.

Codex

tapioca launch codex MODEL

Claude Code

tapioca launch claude MODEL

OpenCode

tapioca launch opencode MODEL

OpenClaw

tapioca launch openclaw MODEL

Hermes

tapioca launch hermes MODEL

06 · Installable knowledge

Give Codex and Claude the playbook.

The repository ships one Agent Skills-compatible package with both Codex and Claude Code manifests. Its use-tapioca skill includes model selection, APIs, media, safety rules, and a structured cross-platform helper.

Browse the plugin →
CODEX

Tapioca Local AI plugin

Add the repository marketplace, then install the local-AI plugin.

codex plugin add tapioca-local-ai@personal
CLAUDE CODE

/tapioca-local-ai:use-tapioca

Test directly from a checkout, then distribute through a Claude marketplace.

claude --plugin-dir ./plugins/tapioca-local-ai

07 · Non-negotiables

Local power needs clear boundaries.

⌁

Loopback first

Never expose an unauthenticated model server by accident.

↓

Disclose downloads

Report size and memory before pulling a very large model.

✓

Permission tools

Treat every model-produced call as an untrusted proposal.

≈

Consent for voices

Possessing an audio file does not imply permission to clone it.

§

License gates

Require the user to review and accept gated model terms themselves.