Tapioca for complete beginners

Local AI,
one step at a time.

You do not need to know what CUDA, GGUF, MLX, diffusion, or a LoRA means. Start with your computer and the thing you want to make. This guide explains every click, command, download, and file.

1
InstallOne app, no account
2
ChooseMatch your memory
3
CreateChat, voice, image, video
4
CustomizeAdd the right LoRA

Lesson 01

Install Tapioca

For most people, the desktop app is the easiest choice. It includes the same engine as the command line but gives you buttons, forms, progress, and a local media gallery.

⌘ Apple Silicon

Install on a Mac

  1. Click the download button.
  2. Open the downloaded .dmg.
  3. Drag Tapioca into Applications.
  4. Open Tapioca and choose Open if macOS asks for confirmation.
Download for Mac
I prefer Terminal
curl -fsSL https://tapioca.rootfruit.cc/install.sh | sh

The installer verifies the download, installs under your user account, and adds Tapioca to new terminal windows.

▦ Windows x64

Install on Windows

  1. Click the download button.
  2. Open the downloaded .exe.
  3. Choose Install for me.
  4. Allow Windows to finish, then open Tapioca.
Download for Windows
I prefer PowerShell
irm https://tapioca.rootfruit.cc/install.ps1 | iex

The installer verifies the archive and adds Tapioca to your user PATH automatically.

What should happen

Open the desktop app and look for a green Tapioca runtime indicator. In a terminal, tapioca version should print a version number.

Before downloading modelsKeep 10–20 GiB of free disk space for beginner text or image models.Video models can require 20–50 GiB.Model files live in ~/.tapioca or %USERPROFILE%\.tapioca.

Lesson 02

Keep Tapioca current without losing anything

Tapioca separates small model-catalog updates from larger software updates. Neither one deletes downloaded models, imported LoRAs, cloned voices, or generated files.

1

Refresh model recipes

tapioca catalog update

Use this when a newly documented model does not appear. It downloads and validates a small checksummed catalog. Desktop does this automatically at startup and also offers Settings → Refresh catalog.

2

Check for a new app version

tapioca update --check

This only checks GitHub Releases; it does not change the computer. Desktop checks automatically and displays an update banner.

3

Install the verified update

tapioca update

The CLI downloads the matching Mac, Windows, or Linux bundle, verifies its SHA-256 checksum, and replaces the executable and runtime. Desktop users click Update now.

Which update do I need?

If the model uses a runtime Tapioca already supports, catalog update is enough. If Tapioca says the model requires a newer runtime or command, install the software update.

Lesson 03

Choose by memory, not hype

A model must fit in memory while it runs. Download size is not the same as memory use, so use Tapioca’s memory recommendation as your first filter.

Your computerSafe first choicesAvoid at first
8–12 GiB memoryqwen3:4b-q4_k_m
chatterbox:nano
Large image and video models
16 GiB memoryqwen3:8b-q4_k_m
FLUX Klein on Mac
SD Turbo on Windows
30B+ LLMs and MiniMax-H3
24–32 GiB memory12B–35B quantized LLMs
SDXL or LTX Video
Models recommending 48 GiB+
48–96 GiB memoryLarge MLX models
Qwen Image Flash
MiniMax-H3 only where licensed
Anything above the catalog recommendation
tapioca catalog update
tapioca catalog

Read each row from left to right: model name, task, download, memory, GPU, platform, and features. If your platform is not listed, choose another variant.

Golden rule

Start small, confirm it works, and move up one size. A smaller responsive model is more useful than a larger model that makes the computer swap or crash.

Lesson 04

Run your first LLM

1

Choose Chat in the app

Select qwen3:4b-q4_k_m for the safest first run on Mac or Windows. It needs about 8 GiB memory.

2

Send a message

Tapioca downloads the model automatically the first time. Keep the app open while the progress bar completes.

tapioca run qwen3:4b-q4_k_m
3

Have a conversation

Ask follow-up questions normally. The model sees the messages in the current conversation.

4

Stop when finished

In Terminal, type /bye or press Ctrl-D. This stops the private model server and releases memory.

What should happen

The first answer may take longer while the model starts. Later messages should begin faster. No prompt or response is uploaded by Tapioca.

The answer is extremely slow

Close memory-heavy apps, choose a smaller model, and avoid context settings larger than you need. Confirm the catalog recommends the model for your memory.

Lesson 05

Clone a voice responsibly

Permission first

Clone only your own voice or a voice whose speaker has clearly agreed. Never use a cloned voice to impersonate someone, bypass verification, or mislead listeners.

Prepare the recording

  • Record one person for 3–10 seconds.
  • Use a quiet room with no music or echo.
  • Speak naturally and save WAV when possible.
  • Write the exact words spoken in the sample.

Save the voice

tapioca voice create narrator \
  --model chatterbox:nano \
  --audio ./narrator.wav \
  --transcript-file ./narrator.txt

narrator is your private local nickname. The audio is copied into Tapioca’s voice folder so moving the original later will not break it.

Generate speech

tapioca tts chatterbox:nano \
  --voice narrator \
  --text "Hello from my local voice." \
  --output hello.wav
What should happen

A playable hello.wav appears in the current folder and in the desktop gallery. The first run takes longer because Tapioca prepares the speech runtime.

Lesson 06

Generate your first image

Open Images, select the model recommended for your hardware, write what should be visible, and choose Generate. Keep the first prompt simple so failures are easy to diagnose.

Mac · 16 GiB+

FLUX.2 Klein

tapioca image flux2-klein:4b-q4-mlx \
  --prompt "A red fox in snow" \
  --output fox.png
Windows · NVIDIA 6 GiB+

SD Turbo CUDA

tapioca image sd-turbo:fp16 \
  --prompt "A red fox in snow" \
  --output fox.png
Windows · AMD or Intel

SD Turbo DirectML

tapioca image sd-turbo:onnx-directml \
  --prompt "A red fox in snow" \
  --output fox.png
96 GiB Mac or NVIDIA 16 GiB+

Krea 2 Turbo

export HF_TOKEN=hf_your_read_token
tapioca pull krea-2-turbo --accept-license
tapioca image krea-2-turbo --prompt "A glass sculpture" --steps 8 --output art.png
Gated model: two approvals

First, sign in at Hugging Face and accept Krea's provider terms. Second, set an approved read token and run --accept-license to record your local Tapioca acknowledgement. That flag does not bypass Hugging Face access. Tapioca never accepts terms for you and does not save the token. Review outputs before sharing them.

Write a useful prompt

Describe subject + setting + lighting + visual style + composition. Example: “A red fox in a snowy pine forest, golden-hour light, detailed wildlife photograph, eye-level portrait.”

Explore, then reproduce

Image, edit, and video use seed 0 by default. Add --random-seed for a new variation and save the number Tapioca prints. Use that number later as --seed NUMBER to repeat the same generation settings. Do not combine the two seed flags.

What should happen

The first run downloads several gigabytes and prepares a private runtime. A progress bar may pause while files load into memory. Later images reuse both downloads.

Lesson 07

Generate motion and video

Video needs much more memory and time than images. Start with a short low-memory clip. Add a starting image when identity or composition matters.

H3 license

MiniMax-H3 is not unrestricted. Its official community license excludes the US, EU, UK, and Republic of Korea; use there requires separate authorization from MiniMax. The publicly downloadable repack and a local acknowledgement do not grant those rights. Check the current terms before downloading or generating.

Mac · 32 GiB+

Wan 2.2

tapioca video wan2.2-video:5b-q8-mlx \
  --prompt "A fox running through snow" \
  --preset low-memory --output fox.mp4
Windows · NVIDIA 8 GiB+

LTX Video

tapioca video ltx-video:2b-fp16 \
  --prompt "A fox running through snow" \
  --preset low-memory --output fox.mp4
48 GiB Mac or 16 GiB NVIDIA

MiniMax-H3 + audio

tapioca video minimax-h3 \
  --image start.png \
  --prompt 'A presenter says exactly: "Hello."' \
  --preset low-memory --output hello.mp4
Choose an approximate duration

Use --seconds 5 instead of calculating a model-compatible frame count. Tapioca prints the selected frame count before generation; the final duration can differ slightly because each video family has its own frame rule. Do not combine --seconds with --frames.

tapioca video minimax-h3 --prompt "A presenter waves" --seconds 5 --output hello.mp4
Long videos

Local video models are best at short shots. Build a 30–60 second video from several 3–5 second clips, then join them. One enormous generation is slower, less stable, and more likely to drift away from the subject.

  • Use --image to anchor the first frame.
  • Use --preset low-memory for the first test.
  • Reduce resolution, frames, or steps when memory runs out.
  • MiniMax-H3 makes native stereo audio; most other video models do not.

Lesson 08

Choose the correct LoRA

A LoRA is a small add-on—not a complete model. It can add a style, subject, motion, or editing behavior only when it was trained for the same base-model architecture you are running.

base model+compatible LoRA+prompt and inputs=output

The six checks to make before downloading

  1. Base model: The model card must name the same family—FLUX Klein, Wan 2.2, MiniMax-H3, SDXL, or another exact architecture.
  2. Task: Image, image editing, or video must match what you are doing.
  3. Runtime: Tapioca supports dynamic LoRAs with MFLUX, CUDA Diffusers, Wan MLX, and MiniMax-H3. ONNX DirectML cannot attach arbitrary LoRAs.
  4. Weight file: Select the exact .safetensors file when a repository contains several.
  5. Inputs: Check how many images are required and their order.
  6. License: Confirm personal or commercial use is permitted for your project.

Inspect before pulling

tapioca adapter inspect hf://OWNER/REPOSITORY
tapioca adapter inspect civitai://MODEL_ID/VERSION_ID
tapioca adapter inspect ms://OWNER/REPOSITORY

Use hf:// for Hugging Face, numeric model/version IDs for Civitai, and ms:// for ModelScope. You can also paste a complete Civitai URL containing modelVersionId.

Select a specific file

tapioca adapter pull hf://OWNER/REPOSITORY \
  --file exact-lora-file.safetensors

Already downloaded it?

tapioca adapter import ~/Downloads/my-lora.safetensors --base minimax-h3 --name my-lora

Import verifies and copies the file into Tapioca. In the desktop app, use Import from computer; Tapioca creates the same managed local:// reference for you.

Apply it gently

tapioca video minimax-h3 \
  --adapter 'hf://OWNER/REPOSITORY#exact-lora-file.safetensors@0.8' \
  --prompt "A cinematic tracking shot" \
  --preset low-memory --output adapted.mp4

The @0.8 is strength. Start around 0.7–0.9. If the output becomes distorted, lower it. Test one LoRA before stacking multiple adapters.

Supported sources and safeguards

  • Hugging Face: Use hf://OWNER/REPOSITORY. Private repositories can use HF_TOKEN.
  • Civitai: Use civitai://MODEL_ID/VERSION_ID or paste the complete version URL. Tapioca rejects checkpoint models when a LoRA is required.
  • ModelScope: Use ms://OWNER/REPOSITORY. modelscope:// is also accepted.
  • Local files: Import regular .safetensors files. Tapioca verifies the file, records its hash and base family, and rejects unsafe paths.
tapioca adapter pull civitai://MODEL_ID/VERSION_ID#adapter.safetensors
tapioca adapter pull ms://OWNER/REPOSITORY#adapter.safetensors

Private sources use environment tokens: HF_TOKEN, CIVITAI_TOKEN, or MODELSCOPE_API_TOKEN. Tapioca never stores these tokens in adapter references or snapshot files.

Verified downloads

Tapioca downloads into a temporary file, checks the provider checksum when available, validates the safetensors header, and only then moves the file into the managed library.

File extension does not prove compatibility

Two files can both end in .safetensors while containing completely different tensor shapes. “It downloads” does not mean “it works with this base model.”

Lesson 09

Reuse downloaded LoRAs

A LoRA only needs to be downloaded or imported once. Tapioca keeps it in its managed adapter library and reuses the local copy whenever you select the same reference.

1. See what is already installed

tapioca adapter list
What should happen

The list shows a reusable reference, its provider, and the exact managed path. Copy the reference from the first column; do not reconstruct it from the filesystem path.

2. Use the reference again

tapioca video minimax-h3   --adapter 'local://my-lora#my-lora.safetensors@0.8'   --prompt "A cinematic tracking shot"   --preset low-memory --output reused.mp4

The same rule applies to an installed hf://, civitai://, or ms:// reference. Tapioca detects the cached file and does not download it again. Changing only @0.8 changes strength; it does not create another copy.

Reuse it in the desktop app

  1. Open Images or Video: Choose a base model that supports LoRAs.
  2. Find LoRA styles: The Installed LoRA menu includes files pulled from providers and files imported from your computer.
  3. Assign LoRA: Choose it, click Assign LoRA, and adjust Strength. You can reorder up to eight adapters.
  4. Generate: Tapioca reuses the managed file. Provider references that are not installed yet are verified and installed automatically.

If the file was downloaded outside Tapioca

tapioca adapter import ~/Downloads/style.safetensors --base minimax-h3 --name style

Use adapter import, not --file, for an existing computer file. The --file option selects a file inside a provider repository. Import copies the weights into the managed library without changing the original.

Move an adapter library to another computer

Copy the complete adapters directory—including each snapshot.json—into the other computer's Tapioca home. The default is ~/.tapioca/adapters on macOS/Linux and %USERPROFILE%\.tapioca\adapters on Windows. Alternatively, import each raw safetensors file again and declare its base model.

Keep metadata together

Do not move only individual managed weight files or rename folders inside the adapter library. The snapshot records provider, checksum, revision, and compatibility information used for safe reuse.

Open the complete import and computer-transfer guide →

Lesson 10

Fix common beginner problems

Nothing seems to happen

Look for download or runtime preparation progress. First runs can take minutes. Keep the app open and confirm free disk space.

The model is not listed

Run tapioca catalog update, then tapioca catalog. Desktop users can choose Refresh catalog in Settings. The verified remote catalog can add recipes for existing runtimes without reinstalling Tapioca; a brand-new runtime still needs tapioca update.

A gated Hugging Face model says access denied

Open its Hugging Face page while signed in, accept the provider terms, create a read token, set HF_TOKEN, and pull once with --accept-license. The token must be present in the same terminal that launches Tapioca.

Windows is not using NVIDIA

Install a current NVIDIA driver, run nvidia-smi, close GPU-heavy apps, and assign Tapioca to the high-performance GPU in Windows Graphics settings.

The computer runs out of memory

Choose a smaller model, use a lower quantization, select low-memory, and reduce video resolution or frames.

The LoRA fails or distorts everything

Recheck the exact base architecture and weight file. Then test the base model alone and retry one LoRA at a lower strength.

You are ready to roll.

Start with one small model and one simple output. The advanced controls will make more sense after the basic loop works once.

Return to Tapioca home →