BigThought docs
BigThought is a proxy between your coding agent and its LLM. You give your agent one URL; it forwards everything to your LLM and scores the model's thinking as it streams.
Connect your client
Wherever these steps say YOUR_PROXY_URL, use the proxy URL shown in your dashboard or settings. Pick the client setting that matches the LLM you chose in BigThought:
| Your LLM in BigThought | Choose this in your client | Base URL |
|---|---|---|
| Ollama or Ollama Cloud | Ollama | YOUR_PROXY_URL |
| OpenAI, OpenRouter, DeepSeek, custom | OpenAI Compatible | YOUR_PROXY_URL/v1 |
| Anthropic | Anthropic, with a custom base URL | YOUR_PROXY_URL |
API key: if you saved one in BigThought settings, type anything into your client's key box (some clients won't save an empty one). Otherwise put your real provider key in the client; BigThought passes it through.
Using an API key instead of the URL
Clients with an API key box can use a BigThought API key (Settings → API keys) instead of the secret URL. Set the base URL to https://bigthought.ai/proxy/v1 (OpenAI-style) or https://bigthought.ai/proxy (Anthropic- or Ollama-style), and put the bt_… key in the client's API key field. Save your LLM provider's key in BigThought's Settings first; BigThought forwards that one.
curl https://bigthought.ai/proxy/v1/chat/completions \
-H "Authorization: Bearer bt_your_key" -H "content-type: application/json" \
-d '{"model": "deepseek-reasoner", "messages": [{"role": "user", "content": "Is 91 prime?"}]}'
Cline (VS Code)
- Open Cline, click the ⚙ settings icon.
- API Provider:
Ollama(orOpenAI Compatible, see the table). - Tick Use custom base URL and paste
YOUR_PROXY_URL(add/v1for OpenAI Compatible). - Model: the model name, e.g.
glm-5.3-flash:cloudordeepseek-r1. For OpenAI Compatible, also fill in Model ID and API Key. - Click Done, ask Cline something, and watch your dashboard.
Roo Code and Kilo Code
Same as Cline: Settings → Providers, choose Ollama or OpenAI Compatible, and paste the base URL from the table.
Cursor
- Cursor Settings → Models.
- Under API Keys, enter a key in OpenAI API Key and switch on Override OpenAI Base URL.
- Base URL:
YOUR_PROXY_URL/v1, then Verify. - Click + Add model and type your model's exact name, then select it in chat.
Cursor needs the hosted service or a public URL. Cursor sends custom-URL requests from its own servers, so a localhost proxy URL won't work. Use the hosted BigThought, or expose your local one with a tunnel.
Continue (VS Code / JetBrains)
Add a model to ~/.continue/config.yaml:
models:
- name: My model via BigThought
provider: ollama # or: openai
model: glm-5.3-flash:cloud
apiBase: YOUR_PROXY_URL # add /v1 when provider is openai
roles: [chat, edit]
Zed
In settings.json:
"language_models": {
"ollama": { "api_url": "YOUR_PROXY_URL" }
}
For OpenAI-style models use "openai": { "api_url": "YOUR_PROXY_URL/v1" }.
Open WebUI
Admin Panel → Settings → Connections: set the Ollama API URL to YOUR_PROXY_URL, or add an OpenAI API connection with YOUR_PROXY_URL/v1.
Aider
OLLAMA_API_BASE=YOUR_PROXY_URL aider --model ollama_chat/glm-5.3-flash:cloud # or OPENAI_API_BASE=YOUR_PROXY_URL/v1 aider --model openai/your-model
Test it with curl
curl YOUR_PROXY_URL/api/chat -d '{
"model": "glm-5.3-flash:cloud",
"messages": [{"role": "user", "content": "Is 91 prime? Think it through."}]
}'
Install with Docker
Self-hosting is free and runs on your own machine: no account, no billing, and your traffic never leaves it. You need Docker.
Everything in one go (recommended)
Save this as docker-compose.yml and run docker compose up -d. It starts BigThought and the Laya scorer.
services:
bigthought:
image: acarli/bigthought:latest
ports: ["127.0.0.1:7860:7860"]
environment:
LAYA_URL: http://laya:8000
DEFAULT_UPSTREAM: http://host.docker.internal:11434
extra_hosts: ["host.docker.internal:host-gateway"]
volumes: [bigthought-data:/data]
laya:
image: acarli/bigthought-laya:latest
volumes: [laya-models:/models]
volumes:
bigthought-data:
laya-models:
Then open http://localhost:7860/app/. Your proxy URL is shown at the top of the dashboard and printed in the logs (docker compose logs bigthought).
The first start downloads the Laya model (about 1.7 GB). On a CPU, scoring takes a few seconds per thought; with an NVIDIA GPU it is about 70 ms. To use a GPU, add deploy: {resources: {reservations: {devices: [{capabilities: [gpu]}]}}} to the laya service (needs the NVIDIA Container Toolkit).
Just the proxy
If you already run laya-serve, or use TypeSafe Jev:
docker run -d --name bigthought -p 127.0.0.1:7860:7860 \ -v bigthought-data:/data --add-host=host.docker.internal:host-gateway \ -e LAYA_URL=http://host.docker.internal:8000 \ acarli/bigthought:latest
For Jev, set LAYA_URL to TypeSafe's API URL and LAYA_API_KEY to your key.
Using Ollama on the same machine
Ollama only listens on 127.0.0.1 by default, which a container can't reach. Either start Ollama with OLLAMA_HOST=0.0.0.0, or on Linux run BigThought with --network host, set DEFAULT_UPSTREAM=http://127.0.0.1:11434, and drop the -p option.
Updating
docker compose pull && docker compose up -d
Which models work
BigThought can only score thinking that the model actually sends:
- Works: Ollama thinking models (GLM, Qwen 3, DeepSeek-R1, gpt-oss…), DeepSeek's reasoner, OpenRouter models that return reasoning, Anthropic Claude with extended thinking, and any model that writes
<think>tags. - Doesn't show thinking: OpenAI's o-series and GPT-5 keep their reasoning hidden, and non-reasoning models have nothing to score. Requests still work normally.
For Ollama, BigThought turns thinking on automatically when your client doesn't say.
Troubleshooting
| Symptom | Fix |
|---|---|
| "choose your LLM first" | Open Settings and pick a provider. |
| "unknown proxy URL" | The URL was regenerated or mistyped. Copy it again from the dashboard. |
| "could not reach your LLM" | Check the base URL in Settings and press Test connection. For local Ollama in Docker, see above. |
| 401 from the LLM | Wrong or missing API key: save it in Settings, or put it in your client. |
| Dashboard shows the call but no thoughts | The model didn't send thinking. See which models work. |
| Thoughts appear but say "not scored" | The scorer is down, or you've used your monthly allowance. |
Privacy
Requests and responses pass straight through to your LLM. The model's thinking is held in memory for your live dashboard and is only saved if you turn on "Keep full thinking" in Session history (encrypted, deleted after 30 days). Each call's summary (time, model, scores, no thinking text) is kept for 12 months so you can see your history. Each thought is sent to the scoring model (Laya, or Jev on the hosted service). API keys you save are encrypted with AES-256-GCM.