Use MiniMax M3 with Continue — velqa.dev
Continue is an open-source VS Code and JetBrains extension for chat, inline editing and agent workflows against the model of your choice. Pointing it at velqa.dev takes two values — a base URL and a key — and gives you one key for chat, editing, agent mode and codebase embeddings. Checkout is Stripe, in USD, by international card.
This page is the full setup: the configuration that works, the two options most people forget, which model to pick, what does *not* work (autocomplete), and how to read the errors you will actually see.
Prerequisites
- VS Code with the Continue extension, or the JetBrains plugin — the configuration file is the same on both.
- A velqa.dev API key. Create an account and generate one from the dashboard; the quickstart does it in three steps and costs nothing.
Configuration
Continue reads ~/.continue/config.yaml. Add a block under models:
name: My Config
version: 0.0.1
schema: v1
models:
- name: MiniMax M3 (Velqa)
provider: openai
model: minimax-m3
apiBase: https://api.velqa.dev/v1
apiKey: sk-...
defaultCompletionOptions:
contextLength: 131072
maxTokens: 16384
capabilities:
- tool_use
roles:
- chat
- edit
- applyprovider: openai does not mean the request goes to OpenAI. It selects Continue's OpenAI-compatible client, which then posts to whatever apiBase says. Velqa speaks that protocol, so nothing else changes.
Three details in that block are worth understanding, because each one causes a distinct failure when it is missing.
apiBase must end in /v1
https://api.velqa.dev/v1, not https://api.velqa.dev. Continue appends /chat/completions to whatever you give it. Drop the /v1 and every request 404s with no useful message.
contextLength and maxTokens are not optional in practice
They cap the context Continue is willing to send and make it truncate or compact *before* the request is too large. Without them, a long chat grows unbounded until the gateway rejects it with:
Trop de tokens demandes. Reessaie plus tard.That error is a per-minute token ceiling, not a broken key — see rate limits. Setting the two values keeps sessions inside the window and, incidentally, spends less of your quota: an oversized prompt is billed whether or not the answer is useful.
Use the model's real window, from the table below.
capabilities: [tool_use] is what turns on agent mode
Continue normally detects tool support from the model name. It does not recognise Velqa's model IDs, so it assumes no tool support and agent mode silently degrades to plain chat — no file reads, no edits applied, no terminal. Declaring tool_use explicitly fixes it. Add image_input as well if you picked a model that accepts images.
Which model to choose
| Model ID | contextLength | maxTokens | Good for | Minimum plan |
|---|---|---|---|---|
hy3 | 262144 | 16384 | General-purpose and agentic, excellent perf/price ratio | Starter |
glm-4.7 | 131072 | 16384 | General chat, strong in French and Arabic | Starter |
glm-4.7-flash | 200000 | 16384 | Fast and cheap, large window | Starter |
deepseek-v4-flash | 65536 | 8192 | High volume, light review — note the smaller window | Starter |
minimax-m3 | 131072 | 16384 | Agentic reasoning, best perf/price ratio | Dev |
kimi-k2.6 | 131072 | 16384 | Agentic coding, long sessions | Dev |
mimo-v2.5 | 262144 | 16384 | General text and agentic, very large window | Dev |
glm-5.2 | 131072 | 16384 | Newer generation, general chat | Dev |
deepseek-v4-pro | 160000 | 16384 | Heavy reasoning flagship | Pro |
qwen3.7-max | 256000 | 16384 | Premium reasoning, 256k context | Pro |
maxTokens is the *output* cap, and 16384 is the ceiling the gateway accepts. deepseek-v4-flash clamps at 8192 upstream — asking for more there is rejected, not truncated.
If you are unsure, start with hy3 on any plan or minimax-m3 on Dev. The current catalogue and USD pricing live at velqa.dev/models and in available models.
Autocomplete: read this before configuring it
Do not give a Velqa model the `autocomplete` role. Continue's autocomplete expects a fill-in-the-middle (FIM) model — one trained to predict the code between a prefix and a suffix. That is a different training format, not a smaller chat model, and Continue's own documentation is explicit that small FIM models beat much larger chat models at the task.
Velqa serves instruction-tuned chat models. Put one in the autocomplete role and you get slow, badly formatted ghost text and a quota bill for every keystroke pause. Leave the role off; keep chat, edit and apply, which is where these models are strong. If you want inline completion, run a local FIM model through Ollama alongside Velqa — Continue happily mixes providers in the same file.
This is the one place where a Velqa-backed Continue setup is genuinely narrower than a Copilot-style setup, and it is better to know it now than after a day of bad suggestions.
Indexing your codebase with @codebase
Continue's @codebase and @docs providers need an embedding model. Velqa serves two multilingual ones on the same key, both included on every plan, so you do not need a second account:
- name: BGE-M3 (Velqa)
provider: openai
model: bge-m3
apiBase: https://api.velqa.dev/v1
apiKey: sk-...
roles:
- embedbge-m3 is the safe default. qwen3-embedding-8b scores higher on multilingual retrieval and takes a larger input; swap the model line to try it. Embeddings are billed on input tokens only — there is no output — so indexing a repository once is cheap. See embeddings.
One honest limitation: Continue's rerank role ships providers for Cohere and Voyage, and Velqa's reranker is not one of them. Leave useReranking off in Continue. Velqa's qwen3-reranker-8b is still available on the API if you are building your own retrieval pipeline — see reranking.
Legacy config.json
Older installations use ~/.continue/config.json. It still loads, but it has no roles and no capabilities, so you lose agent mode:
{
"models": [
{
"name": "MiniMax M3 (Velqa)",
"provider": "openai",
"model": "minimax-m3",
"apiBase": "https://api.velqa.dev/v1",
"apiKey": "sk-..."
}
]
}Migrate to YAML for anything new.
Check that it works
Save the file, open the Continue panel, pick MiniMax M3 (Velqa) in the model selector and send a message. A reply means the connection is good.
If nothing comes back, take Continue out of the equation first:
curl https://api.velqa.dev/v1/chat/completions \
-H "Authorization: Bearer sk-..." \
-H "Content-Type: application/json" \
-d '{"model":"minimax-m3","messages":[{"role":"user","content":"ping"}]}'A reply here with no reply in Continue means the problem is in config.yaml. No reply here either means the key, the model ID or the plan — read on.
Troubleshooting
`401` / invalid key. The full key, no surrounding spaces or quotes, and not revoked in the dashboard. Keys are shown once at creation; generate a new one rather than guessing.
`key not allowed to access model`. Either the ID is misspelled — minimax-m-3 instead of minimax-m3 is the classic — or the model is above your plan. Check it against the table above. A key created before a model was added may also predate its access; regenerate it.
`Trop de tokens demandes. Reessaie plus tard.` A per-minute token ceiling, hit by an oversized request or several at once. Set contextLength and maxTokens as above; if it persists, the session context has grown past what your plan allows per minute. Rate limits has the per-plan figures.
Agent mode does nothing / no tools are called. capabilities: [tool_use] is missing. See above.
`404` on every request. apiBase is missing its /v1.
Requests hang or time out on long answers. Streaming is supported and enabled by default; if a proxy or corporate VPN buffers responses, a long generation can look frozen. Test with the curl above, which streams nothing, to separate the two cases.
Other error codes and their exact meanings are in errors.
Pricing and payment availability
Chat and edit are billed on tokens, embeddings on input tokens only. The current USD catalogue is at velqa.dev/models, and what each plan includes is in plans. Checkout is Stripe, in USD, by international card. Moroccan-issued cards are not accepted here: customers in Morocco are served by velqa.ma, where dirham checkout by Moroccan card opens soon.
See also
- Integrations overview — the other supported tools
- Quickstart — account, key, first request
- Authentication — headers and key handling
- Available models — full catalogue, windows and modalities
