Skip to main content
Velqa

Use MiniMax M3 with Continue — velqa.dev

Continue is an open-source VS Code and JetBrains extension for chat, inline editing and agent workflows against the model of your choice. Pointing it at velqa.dev takes two values — a base URL and a key — and gives you one key for chat, editing, agent mode and codebase embeddings. Checkout is Stripe, in USD, by international card.

This page is the full setup: the configuration that works, the two options most people forget, which model to pick, what does *not* work (autocomplete), and how to read the errors you will actually see.

Prerequisites

  • VS Code with the Continue extension, or the JetBrains plugin — the configuration file is the same on both.
  • A velqa.dev API key. Create an account and generate one from the dashboard; the quickstart does it in three steps and costs nothing.

Configuration

Continue reads ~/.continue/config.yaml. Add a block under models:

name: My Config
version: 0.0.1
schema: v1

models:
  - name: MiniMax M3 (Velqa)
    provider: openai
    model: minimax-m3
    apiBase: https://api.velqa.dev/v1
    apiKey: sk-...
    defaultCompletionOptions:
      contextLength: 131072
      maxTokens: 16384
    capabilities:
      - tool_use
    roles:
      - chat
      - edit
      - apply

provider: openai does not mean the request goes to OpenAI. It selects Continue's OpenAI-compatible client, which then posts to whatever apiBase says. Velqa speaks that protocol, so nothing else changes.

Three details in that block are worth understanding, because each one causes a distinct failure when it is missing.

apiBase must end in /v1

https://api.velqa.dev/v1, not https://api.velqa.dev. Continue appends /chat/completions to whatever you give it. Drop the /v1 and every request 404s with no useful message.

contextLength and maxTokens are not optional in practice

They cap the context Continue is willing to send and make it truncate or compact *before* the request is too large. Without them, a long chat grows unbounded until the gateway rejects it with:

Trop de tokens demandes. Reessaie plus tard.

That error is a per-minute token ceiling, not a broken key — see rate limits. Setting the two values keeps sessions inside the window and, incidentally, spends less of your quota: an oversized prompt is billed whether or not the answer is useful.

Use the model's real window, from the table below.

capabilities: [tool_use] is what turns on agent mode

Continue normally detects tool support from the model name. It does not recognise Velqa's model IDs, so it assumes no tool support and agent mode silently degrades to plain chat — no file reads, no edits applied, no terminal. Declaring tool_use explicitly fixes it. Add image_input as well if you picked a model that accepts images.

Which model to choose

Model IDcontextLengthmaxTokensGood forMinimum plan
hy326214416384General-purpose and agentic, excellent perf/price ratioStarter
glm-4.713107216384General chat, strong in French and ArabicStarter
glm-4.7-flash20000016384Fast and cheap, large windowStarter
deepseek-v4-flash655368192High volume, light review — note the smaller windowStarter
minimax-m313107216384Agentic reasoning, best perf/price ratioDev
kimi-k2.613107216384Agentic coding, long sessionsDev
mimo-v2.526214416384General text and agentic, very large windowDev
glm-5.213107216384Newer generation, general chatDev
deepseek-v4-pro16000016384Heavy reasoning flagshipPro
qwen3.7-max25600016384Premium reasoning, 256k contextPro

maxTokens is the *output* cap, and 16384 is the ceiling the gateway accepts. deepseek-v4-flash clamps at 8192 upstream — asking for more there is rejected, not truncated.

If you are unsure, start with hy3 on any plan or minimax-m3 on Dev. The current catalogue and USD pricing live at velqa.dev/models and in available models.

Autocomplete: read this before configuring it

Do not give a Velqa model the `autocomplete` role. Continue's autocomplete expects a fill-in-the-middle (FIM) model — one trained to predict the code between a prefix and a suffix. That is a different training format, not a smaller chat model, and Continue's own documentation is explicit that small FIM models beat much larger chat models at the task.

Velqa serves instruction-tuned chat models. Put one in the autocomplete role and you get slow, badly formatted ghost text and a quota bill for every keystroke pause. Leave the role off; keep chat, edit and apply, which is where these models are strong. If you want inline completion, run a local FIM model through Ollama alongside Velqa — Continue happily mixes providers in the same file.

This is the one place where a Velqa-backed Continue setup is genuinely narrower than a Copilot-style setup, and it is better to know it now than after a day of bad suggestions.

Indexing your codebase with @codebase

Continue's @codebase and @docs providers need an embedding model. Velqa serves two multilingual ones on the same key, both included on every plan, so you do not need a second account:

  - name: BGE-M3 (Velqa)
    provider: openai
    model: bge-m3
    apiBase: https://api.velqa.dev/v1
    apiKey: sk-...
    roles:
      - embed

bge-m3 is the safe default. qwen3-embedding-8b scores higher on multilingual retrieval and takes a larger input; swap the model line to try it. Embeddings are billed on input tokens only — there is no output — so indexing a repository once is cheap. See embeddings.

One honest limitation: Continue's rerank role ships providers for Cohere and Voyage, and Velqa's reranker is not one of them. Leave useReranking off in Continue. Velqa's qwen3-reranker-8b is still available on the API if you are building your own retrieval pipeline — see reranking.

Legacy config.json

Older installations use ~/.continue/config.json. It still loads, but it has no roles and no capabilities, so you lose agent mode:

{
  "models": [
    {
      "name": "MiniMax M3 (Velqa)",
      "provider": "openai",
      "model": "minimax-m3",
      "apiBase": "https://api.velqa.dev/v1",
      "apiKey": "sk-..."
    }
  ]
}

Migrate to YAML for anything new.

Check that it works

Save the file, open the Continue panel, pick MiniMax M3 (Velqa) in the model selector and send a message. A reply means the connection is good.

If nothing comes back, take Continue out of the equation first:

curl https://api.velqa.dev/v1/chat/completions \
  -H "Authorization: Bearer sk-..." \
  -H "Content-Type: application/json" \
  -d '{"model":"minimax-m3","messages":[{"role":"user","content":"ping"}]}'

A reply here with no reply in Continue means the problem is in config.yaml. No reply here either means the key, the model ID or the plan — read on.

Troubleshooting

`401` / invalid key. The full key, no surrounding spaces or quotes, and not revoked in the dashboard. Keys are shown once at creation; generate a new one rather than guessing.

`key not allowed to access model`. Either the ID is misspelled — minimax-m-3 instead of minimax-m3 is the classic — or the model is above your plan. Check it against the table above. A key created before a model was added may also predate its access; regenerate it.

`Trop de tokens demandes. Reessaie plus tard.` A per-minute token ceiling, hit by an oversized request or several at once. Set contextLength and maxTokens as above; if it persists, the session context has grown past what your plan allows per minute. Rate limits has the per-plan figures.

Agent mode does nothing / no tools are called. capabilities: [tool_use] is missing. See above.

`404` on every request. apiBase is missing its /v1.

Requests hang or time out on long answers. Streaming is supported and enabled by default; if a proxy or corporate VPN buffers responses, a long generation can look frozen. Test with the curl above, which streams nothing, to separate the two cases.

Other error codes and their exact meanings are in errors.

Pricing and payment availability

Chat and edit are billed on tokens, embeddings on input tokens only. The current USD catalogue is at velqa.dev/models, and what each plan includes is in plans. Checkout is Stripe, in USD, by international card. Moroccan-issued cards are not accepted here: customers in Morocco are served by velqa.ma, where dirham checkout by Moroccan card opens soon.

See also