BigThought docs

BigThought is a proxy between your coding agent and its LLM. You give your agent one URL; it forwards everything to your LLM and scores the model's thinking as it streams.

Connect your client Install with Docker Which models work Troubleshooting Privacy

Connect your client

Wherever these steps say YOUR_PROXY_URL, use the proxy URL shown in your dashboard or settings. Pick the client setting that matches the LLM you chose in BigThought:

Your LLM in BigThoughtChoose this in your clientBase URL
Ollama or Ollama CloudOllamaYOUR_PROXY_URL
OpenAI, OpenRouter, DeepSeek, customOpenAI CompatibleYOUR_PROXY_URL/v1
AnthropicAnthropic, with a custom base URLYOUR_PROXY_URL

API key: if you saved one in BigThought settings, type anything into your client's key box (some clients won't save an empty one). Otherwise put your real provider key in the client; BigThought passes it through.

Using an API key instead of the URL

Clients with an API key box can use a BigThought API key (Settings → API keys) instead of the secret URL. Set the base URL to https://bigthought.ai/proxy/v1 (OpenAI-style) or https://bigthought.ai/proxy (Anthropic- or Ollama-style), and put the bt_… key in the client's API key field. Save your LLM provider's key in BigThought's Settings first; BigThought forwards that one.

curl https://bigthought.ai/proxy/v1/chat/completions \
  -H "Authorization: Bearer bt_your_key" -H "content-type: application/json" \
  -d '{"model": "deepseek-reasoner", "messages": [{"role": "user", "content": "Is 91 prime?"}]}'

Cline (VS Code)

  1. Open Cline, click the ⚙ settings icon.
  2. API Provider: Ollama (or OpenAI Compatible, see the table).
  3. Tick Use custom base URL and paste YOUR_PROXY_URL (add /v1 for OpenAI Compatible).
  4. Model: the model name, e.g. glm-5.3-flash:cloud or deepseek-r1. For OpenAI Compatible, also fill in Model ID and API Key.
  5. Click Done, ask Cline something, and watch your dashboard.

Roo Code and Kilo Code

Same as Cline: Settings → Providers, choose Ollama or OpenAI Compatible, and paste the base URL from the table.

Cursor

  1. Cursor Settings → Models.
  2. Under API Keys, enter a key in OpenAI API Key and switch on Override OpenAI Base URL.
  3. Base URL: YOUR_PROXY_URL/v1, then Verify.
  4. Click + Add model and type your model's exact name, then select it in chat.

Cursor needs the hosted service or a public URL. Cursor sends custom-URL requests from its own servers, so a localhost proxy URL won't work. Use the hosted BigThought, or expose your local one with a tunnel.

Continue (VS Code / JetBrains)

Add a model to ~/.continue/config.yaml:

models:
  - name: My model via BigThought
    provider: ollama            # or: openai
    model: glm-5.3-flash:cloud
    apiBase: YOUR_PROXY_URL        # add /v1 when provider is openai
    roles: [chat, edit]

Zed

In settings.json:

"language_models": {
  "ollama": { "api_url": "YOUR_PROXY_URL" }
}

For OpenAI-style models use "openai": { "api_url": "YOUR_PROXY_URL/v1" }.

Open WebUI

Admin Panel → Settings → Connections: set the Ollama API URL to YOUR_PROXY_URL, or add an OpenAI API connection with YOUR_PROXY_URL/v1.

Aider

OLLAMA_API_BASE=YOUR_PROXY_URL aider --model ollama_chat/glm-5.3-flash:cloud
# or
OPENAI_API_BASE=YOUR_PROXY_URL/v1 aider --model openai/your-model

Test it with curl

curl YOUR_PROXY_URL/api/chat -d '{
  "model": "glm-5.3-flash:cloud",
  "messages": [{"role": "user", "content": "Is 91 prime? Think it through."}]
}'

Install with Docker

Self-hosting is free and runs on your own machine: no account, no billing, and your traffic never leaves it. You need Docker.

Everything in one go (recommended)

Save this as docker-compose.yml and run docker compose up -d. It starts BigThought and the Laya scorer.

services:
  bigthought:
    image: acarli/bigthought:latest
    ports: ["127.0.0.1:7860:7860"]
    environment:
      LAYA_URL: http://laya:8000
      DEFAULT_UPSTREAM: http://host.docker.internal:11434
    extra_hosts: ["host.docker.internal:host-gateway"]
    volumes: [bigthought-data:/data]
  laya:
    image: acarli/bigthought-laya:latest
    volumes: [laya-models:/models]
volumes:
  bigthought-data:
  laya-models:

Then open http://localhost:7860/app/. Your proxy URL is shown at the top of the dashboard and printed in the logs (docker compose logs bigthought).

The first start downloads the Laya model (about 1.7 GB). On a CPU, scoring takes a few seconds per thought; with an NVIDIA GPU it is about 70 ms. To use a GPU, add deploy: {resources: {reservations: {devices: [{capabilities: [gpu]}]}}} to the laya service (needs the NVIDIA Container Toolkit).

Just the proxy

If you already run laya-serve, or use TypeSafe Jev:

docker run -d --name bigthought -p 127.0.0.1:7860:7860 \
  -v bigthought-data:/data --add-host=host.docker.internal:host-gateway \
  -e LAYA_URL=http://host.docker.internal:8000 \
  acarli/bigthought:latest

For Jev, set LAYA_URL to TypeSafe's API URL and LAYA_API_KEY to your key.

Using Ollama on the same machine

Ollama only listens on 127.0.0.1 by default, which a container can't reach. Either start Ollama with OLLAMA_HOST=0.0.0.0, or on Linux run BigThought with --network host, set DEFAULT_UPSTREAM=http://127.0.0.1:11434, and drop the -p option.

Updating

docker compose pull && docker compose up -d

Which models work

BigThought can only score thinking that the model actually sends:

For Ollama, BigThought turns thinking on automatically when your client doesn't say.

Troubleshooting

SymptomFix
"choose your LLM first"Open Settings and pick a provider.
"unknown proxy URL"The URL was regenerated or mistyped. Copy it again from the dashboard.
"could not reach your LLM"Check the base URL in Settings and press Test connection. For local Ollama in Docker, see above.
401 from the LLMWrong or missing API key: save it in Settings, or put it in your client.
Dashboard shows the call but no thoughtsThe model didn't send thinking. See which models work.
Thoughts appear but say "not scored"The scorer is down, or you've used your monthly allowance.

Privacy

Requests and responses pass straight through to your LLM. The model's thinking is held in memory for your live dashboard and is only saved if you turn on "Keep full thinking" in Session history (encrypted, deleted after 30 days). Each call's summary (time, model, scores, no thinking text) is kept for 12 months so you can see your history. Each thought is sent to the scoring model (Laya, or Jev on the hosted service). API keys you save are encrypted with AES-256-GCM.

Terms · Privacy