Skip to main content
After completing this guide, you’ll have a running Spaceduck gateway with a chat model connected and ready to use.

Install

Prefer to install manually? See the From Source guide.

Configure your chat model

Open http://localhost:3000 and click the Settings icon in the sidebar, then go to Chat.
  1. Select a Provider from the dropdown (e.g., llama.cpp, Bedrock, Gemini)
  2. Enter the Base URL if using a local provider
  3. Set your API key if using a cloud provider
  4. Choose a Model
Changes to provider, model, and system prompt hot-apply immediately — no restart needed.
If you’re unsure which provider to start with, see the Model Providers overview for a comparison.

Send your first message

Go back to the chat view and type something. You should see a streaming response from your configured model.
If you see a response, Spaceduck is working. Your conversations are now being stored in SQLite with automatic fact extraction.

Optional: enable semantic recall

By default, Spaceduck uses keyword search (FTS5) for cross-conversation memory. To enable vector-based semantic recall:
  1. Go to Settings > Memory
  2. Toggle Semantic recall on
  3. Select an embedding provider and model
  4. Click Test to verify the connection
Semantic recall requires a separate embedding model. You can use the same provider as your chat model, or run a dedicated embedding server. See llama.cpp embeddings for a local setup.

Optional: install tools

Enables web browsing and page interaction via the browser tool.
When marker_single is on your PATH, the marker_scan tool is automatically registered. Upload PDFs via the paperclip button in the chat UI.
Web UI: Click the mic button in the chat input. Hold to record, release to transcribe via Whisper or AWS Transcribe.
When whisper is on your PATH, the mic button appears in the chat UI.Alternatively, configure AWS Transcribe credentials in your environment for cloud-based STT.Desktop app (macOS): Press the Fn/Globe key from any application to activate a system-wide dictation pill with real-time waveform. See Desktop App for setup.

Choose your client

Spaceduck supports multiple clients — all connecting to the same gateway.

What you now have

  • A running gateway at http://localhost:3000
  • A chat model connected and streaming responses
  • Persistent conversation history in SQLite
  • Automatic fact extraction after every turn
  • Keyword-based cross-conversation recall (or semantic recall if you enabled it)

Next steps

Model Providers

Learn about chat vs embedding providers and the two-server pattern.

Memory

Understand how Spaceduck remembers and recalls facts across conversations.