GreenPT Docs

Model Cards

Full GreenPT model catalog (chat, coding, embeddings, speech-to-text, and reranking) with pricing per model.

Full GreenPT model catalog. Filter by category and compare pricing. Token prices are in EUR per 1M tokens; speech-to-text models are billed per hour of audio. To list models programmatically, see the Models API.

Catalog

kimi-k3Moonshot AINew

Moonshot Kimi K3, an open-weight 2.8T-parameter native multimodal agentic model with hybrid reasoning for long-horizon coding, knowledge work, and reasoning.

Input€3.30
Cached input€0.825
Output€16.50
Context1M
ReleasedJul 2026
ChatHybrid reasoningAgentic codingVision
glm-5.2z.aiNew

GreenPT Code flagship coding model with strong reasoning, agentic tool-use, and multi-file software engineering performance. See the GreenPT Code guide for IDE setup. Also available as output-compression variants (glm-5.2-caveman, -honey, -ponytail) at the same price.

Input€1.10
Cached input€0.275
Output€4.40
Context1M
Tensor typeBF16 · F32
Released2026
Code generationAgentic coding tasksMulti-file reasoningTool use
deepseek-v4-flash-0731DeepSeekNew

DeepSeek V4 Flash, 2026-07-31 snapshot. A fast, low-cost reasoning model with a 1M-token context window and tool use. Reasoning is always on, so allow enough output budget for it.

Input€0.14
Cached input€0.04
Output€0.35
Context1M
ReleasedAug 2026
ChatReasoningTool useLong context
minimax-m2.5MiniMax

MiniMax reasoning model optimized for agentic coding and long-horizon workflows.

Input€0.33
Output€1.32
Context192k
Max output128k
Released2026
ChatReasoningAgentic codingLong-horizon tasks
kimi-k2.6Moonshot AI

Moonshot Kimi K2.6 reasoning model with strong agentic coding and multimodal input.

Input€0.66
Cached input€0.22
Output€3.75
Context256k
Max output256k
Released2026
ChatReasoningAgentic codingVision
kimi-k2.7-codeMoonshot AI

Moonshot Kimi K2.7 Code, tuned for software engineering and agentic coding tasks.

Input€0.77
Cached input€0.165
Output€3.85
Context256k
Max output256k
Released2026
ChatReasoningCodingAgentic tasks
gemma4GreenPT
Backed by Google Gemma 4

GreenPT-branded next-generation chat model with long-context multimodal reasoning. Recommended default for chat completions.

Input€0.50
Output€1.50
Context256k
Max output32k
Released2026
ChatMultimodal reasoningLong-contextDefault model
green-rGreenPT
Backed by GPT-OSS

GreenPT-branded reasoning model for advanced analysis, writing, and content generation.

Input€0.35
Output€0.95
Context128k
Max output32k
ReleasedAug 2025
TextImagesDocumentsMultilingualAdvanced reasoningWriting & content generation
green-r-rawGreenPT
Backed by GPT-OSS

Direct access to the same reasoning stack as green-r without the GreenPT system prompt.

Input€0.35
Output€0.95
Context128k
Max output32k
ReleasedAug 2025
TextImagesDocumentsMultilingualNo system prompt
green-lGreenPT
Backed by Mistral Small 3.2 24B

GreenPT-branded chat model tuned for multilingual writing, image understanding, and Dutch grammar guardrails.

Input€0.25
Output€0.80
Context128k
Max output32k
ReleasedJun 2025
TextImagesDocumentsMultilingualWriting assistantDutch grammar guardrails
green-l-rawGreenPT
Backed by Mistral Small 3.2 24B

Direct access to the same GreenPT-backed model as green-l without the built-in system prompt.

Input€0.25
Output€0.80
Context128k
Max output32k
ReleasedJun 2025
TextImagesDocumentsMultilingualNo system prompt
qwen3.5-397b-a17bQwen

State-of-the-art Qwen model optimized for code generation, agentic tasks, and logical reasoning.

Input€0.70
Output€4.35
Context250k
Max output16k
ReleasedMar 2026
ChatCode generationAgentic tasksLogical reasoningState-of-the-art (March 2026)
mistral-small-3.2-24b-instruct-2506Mistral

General-purpose model with balanced cost, function calling, and multimodal chat support.

Input€0.20
Output€0.40
Context128k
Max output32k
ReleasedJun 2025
ChatVisionFunction callingMultilingual
gpt-oss-120bOpenAI

Large open model with vision, function calling, and long-context reasoning support.

Input€0.20
Output€0.70
Context128k
Max output32k
ReleasedAug 2025
ChatVisionLong-context reasoningFunction calling
qwen3-235b-a22b-instruct-2507Qwen

High-context multilingual reasoning model with a 250k token window.

Input€0.90
Output€2.70
Context250k
Max output16k
ReleasedJul 2025
ChatLong-contextReasoningMultilingual250k context window
qwen3-coder-30b-a3b-instructQwen

Qwen coding model built for generation, completion, and debugging workflows.

Input€0.25
Output€0.95
Context256k
Max output32k
ReleasedJul 2025
ChatCode generationCompletionDebugging
llama-3.3-70b-instructMeta

Meta instruction-following model with strong multilingual support for general chat use cases.

Input€1.10
Output€1.10
Context128k
Max output16k
ReleasedDec 2024
ChatInstruction followingMultilingual
qwen3.6-35b-a3bQwen

Frontier small Qwen model optimized for agentic and reasoning tasks across many languages.

Input€0.30
Output€1.80
Context256k
Max output32k
Released2026
ChatCodeVisionAgentic & reasoning tasks
mistral-medium-3.5-128bMistral

Unified Mistral model with strong instruct, reasoning, coding, and vision performance.

Input€1.80
Output€9.00
Context180k
Max output16k
Released2026
ChatVisionReasoningCoding
green-embeddingGreenPT
Backed by Qwen3-Embedding-4B

Dense multilingual embedding model for semantic search, retrieval, and RAG pipelines. Supports Matryoshka Representation Learning, so output dimensions are configurable from 32 up to 2560.

Input€0.20
OutputNo output charge
Context32k
Dimensionsup to 2560
ReleasedJun 2025
Dense embeddingsMultilingual (100+ languages)Semantic searchRAGMRL (custom dimensions)Instruction-aware
qwen3-embedding-8bQwen
Backed by Qwen3-Embedding-8B

State-of-the-art multilingual embedding model for semantic search, retrieval, and RAG pipelines. Supports Matryoshka Representation Learning, so output dimensions are configurable from 32 up to 4096.

Input€0.25
OutputNo output charge
Context32k
Dimensionsup to 4096
Released2026
Dense embeddingsMultilingual (100+ languages)Semantic searchRAGMRL (custom dimensions)Instruction-aware
bge-multilingual-gemma2BAAI

Compact multilingual embedding model, a cost-efficient option for large corpora where an 8k input window is enough.

Input€0.05
OutputNo output charge
Context8k
Dimensions3584
Released2024
EmbeddingsMultilingual
green-sGreenPT50% off until 30 Sep

Speech-to-text model for pre-recorded and live transcription.

Input€0.23/hr (€0.12 promo)
Output€0.31/hr live (€0.16 promo)
Released2025
Pre-recorded transcriptionLive transcription50% off until 30 September 2026
green-s-proGreenPTPro50% off until 30 Sep

Advanced speech-to-text model with multilingual transcription options.

Input€0.23/hr recorded (€0.12 promo)
Output€0.31/hr live (€0.16 promo)
Released2025
Pre-recorded transcriptionLive transcriptionMultilingual input €0.28/hr (€0.14 promo)Multilingual live €0.37/hr (€0.19 promo)50% off until 30 September 2026
green-rerankGreenPT
Backed by Qwen3-Reranker-4B

Cross-encoder reranking model that scores and reorders search or RAG results by relevance to a query.

Input€0.12
OutputNo output charge
Context32k
ReleasedJun 2025
Document rerankingRAG result scoringMultilingual (100+ languages)Instruction-aware

Deprecated model ids

A pending id keeps answering and is served by its replacement until the listed date. A removed id is no longer listed by GET /v1/models from that date, and requests to it return 400 Unsupported model. A directGET /v1/models/{id} lookup returns 404.

devstral-2-123b-instruct-2512MistralDeprecation pendingserved by glm-5.2 from 1 September 2026
gemma-3-27b-itGoogleDeprecation pendingserved by gemma4 from 1 September 2026
holo2-30b-a3bH CompanyDeprecation pendingserved by qwen3.6-35b-a3b from 1 September 2026
qwen3-coder-30b-a3b-instructQwenDeprecation pendingserved by deepseek-v4-flash-0731 from 1 October 2026
pixtral-12b-2409MistralDeprecation pendingserved by mistral-small-3.2-24b-instruct-2506 from 1 October 2026
deepseek-r1-distill-llama-70bDeepSeekDeprecatedremoved on 18 August 2026 · use deepseek-v4-flash-0731
devstral-small-2505MistralDeprecatedremoved on 18 August 2026 · use deepseek-v4-flash-0731
llama-3.1-8b-instructMetaDeprecatedremoved on 18 August 2026 · use deepseek-v4-flash-0731
mistral-nemo-instruct-2407MistralDeprecatedremoved on 18 August 2026 · use deepseek-v4-flash-0731
voxtral-small-24b-2507MistralDeprecatedremoved on 18 August 2026 · use green-s-pro

Built-in system prompts

Some ids carry a system prompt of their own, so the id you send decides which rules the model gets on top of yours:

  • Compression models (glm-5.2-caveman, -honey and -ponytail, each also in -lite and -ultra) carry an output-compression ruleset: the same model and price per token as glm-5.2, with a shorter answer. See Compression models.
  • GreenPT models (green-l, green-r) carry GreenPT's Sustainability system prompt, optimized for our use cases; their raw twins (green-l-raw, green-r-raw) carry none, so yours is the only system prompt sent upstream.
  • Every other id in the catalog carries no system prompt.

When a model carries one and you send one too, both go upstream as a single system message: yours first, the model's rules last. Picking the id is an explicit opt-in to that style, so on a conflict the model's rules win over a generic prompt your client happens to send. A developer message counts as a system message here and merges the same way.

Catalog Notes

All models are served through the GreenPT API proxy with an OpenAI-compatible interface. Prices are in EUR per million tokens unless noted otherwise, and API pricing applies to API usage only. For how we measure the environmental footprint of a request, see Carbon calculations.

On this page