Models available on velqa.dev
All the models below are reachable through velqa.dev's OpenAI-compatible API (https://api.velqa.dev/v1/chat/completions) and Anthropic-compatible API (https://api.velqa.dev/v1/messages). Checkout is Stripe, in USD, by international card. Moroccan-issued cards are not accepted here: customers in Morocco are served by velqa.ma, where dirham checkout by Moroccan card opens soon.
Model table
| Model ID | Source provider | Use case | Type |
|---|---|---|---|
glm-4.7-flash | Zhipu AI | Fast chat, everyday tasks | Fast/Chat |
deepseek-v4-flash | DeepSeek | Cheap high-volume, light code review | Fast/Chat |
hy3 | Tencent | Generalist and agentic | Chat/Agentic |
glm-4.7 | Zhipu AI | Chat and agentic coding | Chat/Coding |
minimax-m3 | MiniMax | Agentic reasoning | Coding |
mimo-v2.5 | Xiaomi | Text generalist, chat and agentic | Chat/Agentic |
kimi-k2.6 | Moonshot AI | Agentic coding, long sessions, 131k context | Coding |
glm-5.2 | Zhipu AI | Long autonomous tasks, 128k context | Coding |
deepseek-v4-pro | DeepSeek | Flagship heavy reasoning, 160k context | Premium |
qwen3.7-max | Alibaba | Premium reasoning, 256k context | Premium |
qwen3.8-max | Alibaba | Flagship coding and professional work, 256k context | Premium |
kimi-k3 | Moonshot AI | Heavy agentic reasoning, 1,048,576-token context | Premium |
bge-m3 | BAAI | Multilingual embeddings | Embeddings |
qwen3-embedding-8b | Alibaba | High-quality multilingual embeddings | Embeddings |
qwen3-reranker-8b | Alibaba | Multilingual reranking | Rerank |
whisper-large-v3 | OpenAI (via DeepInfra) | Multilingual audio transcription, max 10 min | Audio (STT) |
kokoro-tts | hexgrad (via DeepInfra) | Text-to-speech, 8 languages, no Arabic | Audio (TTS) |
qwen3-tts | Alibaba (via DeepInfra) | Text-to-speech with Arabic and tone control | Audio (TTS) |
flux-1-schnell | Black Forest Labs (via DeepInfra) | Text-to-image, draft/volume tier | Image |
flux-2-klein-4b | Black Forest Labs (via DeepInfra) | Text-to-image generation | Image |
flux-2-pro | Black Forest Labs (via DeepInfra) | Text-to-image, quality tier | Image |
fastwan-t2v-1.3b | FastVideo (via DeepInfra) | Text-to-video, 480p, seconds not minutes | Video |
wan2.2-t2v-a14b | Wan-AI (via DeepInfra) | Text-to-video generation, 5 seconds | Video |
Availability per plan
| Model | Starter | Dev | Pro | Recharge Boost |
|---|---|---|---|---|
glm-4.7-flash | Yes | Yes | Yes | Yes |
deepseek-v4-flash | Yes | Yes | Yes | Yes |
hy3 | Yes | Yes | Yes | Yes |
glm-4.7 | Yes | Yes | Yes | Yes |
minimax-m3 | No | Yes | Yes | Yes |
mimo-v2.5 | No | Yes | Yes | Yes |
kimi-k2.6 | No | Yes | Yes | Yes |
glm-5.2 | No | Yes | Yes | Yes |
deepseek-v4-pro | No | No | Yes | Yes |
qwen3.7-max | No | No | Yes | Yes |
qwen3.8-max | No | No | No | Yes |
kimi-k3 | No | No | No | Yes |
bge-m3 | Yes | Yes | Yes | Yes |
qwen3-embedding-8b | Yes | Yes | Yes | Yes |
qwen3-reranker-8b | Yes | Yes | Yes | Yes |
whisper-large-v3 | No | No | No | Yes |
kokoro-tts | No | No | No | Yes |
qwen3-tts | No | No | No | Yes |
flux-1-schnell | No | No | No | Yes |
flux-2-klein-4b | No | No | No | Yes |
flux-2-pro | No | No | No | Yes |
fastwan-t2v-1.3b | No | No | No | Yes |
wan2.2-t2v-a14b | No | No | No | Yes |
Boost-only models
kimi-k3 and qwen3.8-max were added on 2026-08-09 and are the first chat models sold outside every subscription. Their per-token cost is beyond what a fixed-price plan can absorb — kimi-k3 alone costs several times an included model's budget for the same turn — so they are sold per token against a prepaid Boost balance instead of being folded into a tier and paid for by everyone else's quota.
kimi-k3: heavy agentic reasoning over a 1,048,576-token window. Generation is slow (single-digit tokens per second), so it suits work where reasoning quality outweighs latency, not interactive chat.qwen3.8-max: generalist flagship aimed at code and professional work, 256k context on the primary route.
Calling either one with a subscription key that has no Boost balance is rejected: they are not in that key's allowed-model list. Per-token prices for both are on the pricing page.
Media models added on 2026-08-09
Four at once, all on DeepInfra like the rest of the media catalogue — no new provider relationship.
flux-1-schnell($0.002/image): 15x cheaper than the default tier. This is the drafting model, for the runs where paying $0.03 an attempt makes no sense.flux-2-pro($0.032/image): clearly better thanflux-2-klein-4bfor ~7% more.flux-2-klein-9bcosts exactly the same and adds nothing, so it was not added.qwen3-tts($44/M characters): fills the gapkokoro-ttsleaves — kokoro speaks 8 languages, none of them Arabic. Also adds forced language and a natural-language tone directive. At 22x kokoro's price it is for the cases kokoro cannot cover, not the default.fastwan-t2v-1.3b($0.005/second): 30x cheaper thanwan2.2-t2v-a14band, more to the point, ~100x faster — a measured 2.5 seconds of generation against 4m35s for the same 5-second clip. It renders at 480p with less detail: this is the iteration tier,wan2.2stays the quality tier.
Notes
- Model identifiers (
model_id) are stable — these are the ones to use in your tool configs. - The Starter plan is designed for students and exploration; advanced coding models require the Dev plan at minimum.
- Recharge Boost mode (pay-as-you-go) gives access to every model from a prepaid credit balance, with no subscription.
- Current USD pricing and checkout availability are shown in the dashboard; customers in Morocco are served by velqa.ma, where dirham checkout by Moroccan card opens soon.
