Skip to content

Model Selection

Simply select “auto” in the OpenAI model selection to activate CompanyGPT’s dynamic routing. The system analyzes your prompt and automatically chooses the most efficient OpenAI model: fast, smaller models for standard questions and high-end models for complex analyses. This saves time and token costs without any manual effort.

  • Fast & inexpensive → Mini / Flash / Nano / Haiku
  • Standard & reliable → large all-round models
  • Complex & critical → most powerful models
  • EU / internal / data protection → STACKIT models

  • For: complex problems and maximum intelligence
  • When: demanding coding, scientific analyses, strategic planning
  • Why: OpenAI’s most advanced model with unmatched logical reasoning
  • For: deep thinking with high reliability
  • When: complex documents, data analyses, advanced text work
  • Why: strong GPT-5 generation performance optimized for practical use
  • For: high quality at high speed
  • When: everyday programming tasks, structured research, logical iterations
  • Why: excellent balance of GPT-5 intelligence and fast processing time
  • For: ultra-fast answers and simple assistance
  • When: quick questions, simple formatting, text corrections
  • Why: extremely resource-efficient model of the GPT-5 family with minimal latency
  • For: reliable problem-solving and standard contexts
  • When: text optimization, research, logical comparisons
  • Why: the proven, refined workhorse for demanding everyday tasks
  • For: fast processing at low cost
  • When: structured text creation, data pre-sorting, simple conversations
  • Why: compact version of GPT-4.1, optimized for efficiency and speed
  • For: versatile multimodality and fluid interaction
  • When: image and speech processing, creative drafts, brainstorming
  • Why: the established flagship model for fast, multimedia tasks
  • For: high speed at minimal cost
  • When: simple chat assistants, fast filtering of large data volumes
  • Why: extremely inexpensive lightweight model with solid base intelligence
  • For: automated efficiency without manual selection
  • When: varying tasks, straightforward workflows
  • Why: automatically selects the most suitable model based on the complexity of your request
  • For: image generation
  • When: when images need to be generated
  • Why: OpenAI’s image generation model

  • For: maximum speed and cost efficiency
  • When: quick follow-up questions, simple bulk tasks, real-time translations
  • Why: extremely short response time at unbeatably low cost
  • For: autonomous reasoning at high speed
  • When: complex workflows, deep analyses, demanding coding
  • Why: Google’s fast model with flexible thinking levels for smarter tasks
  • For: image analysis, image generation, image editing
  • When: text-to-image generation, image editing with prompts (image + text), and composition of multiple images
  • Why: Google’s image models integrated into CompanyGPT

  • For: programming, complex text processing, and demanding all-round tasks (Recommended)
  • When: software development, code refactoring, deep text comprehension
  • Why: the sweet spot of the series. Opus-class performance at Sonnet pricing (1M context)
  • For: very fast processing with high logical precision
  • When: filtering large data volumes, UI-based chatbots, simple to medium tasks at scale
  • Why: very fast and cost-efficient (less reasoning than Sonnet/Opus)

These open-source models run in the STACKIT Cloud (EU/Germany) and are particularly suited for workloads with high requirements for data protection, data sovereignty, and internal compliance.

  • For: highest OSS quality for complex text tasks
  • When: when internally hosted top performance is required instead of maximum speed
  • Why: large open-source model for strong analytical and linguistic results
  • For: multimodal high-end analysis with image and text understanding
  • When: visual document analysis, complex image-text tasks, demanding inference
  • Why: very powerful VL model for deeper understanding of multimodal content

cortecs/Llama-3.3-70B-Instruct-FP8-Dynamic

Section titled “cortecs/Llama-3.3-70B-Instruct-FP8-Dynamic”
  • For: more demanding generation and reasoning in the EU stack
  • When: more complex enterprise questions, longer responses, better level of detail
  • Why: the 70B class provides significantly more quality than small models while retaining OSS flexibility
  • For: versatile instruct tasks with good efficiency
  • When: internal assistants, knowledge work, structured text production
  • Why: strong mid-range between small fast and large expensive OSS models
  • For: versatile text tasks and strong multilingual understanding
  • When: creative writing, logical problem-solving, precise translations
  • Why: the latest Qwen generation with an optimal balance of compact size and high intelligence
  • For: lightweight open-source text tasks
  • When: cost-sensitive internal workflows with controllable infrastructure
  • Why: compact OSS approach for solid quality with lower resource requirements
  • For: multimodal embeddings (text/image) for search and retrieval
  • When: semantic search, RAG indexing, similarity search across mixed data
  • Why: specialized in vector representations rather than classic chat responses
  • For: high-quality text embeddings for retrieval and ranking
  • When: vector search, document retrieval, relevance ranking in RAG pipelines
  • Why: proven embedding model for precise semantic search applications

  • “I just want a very good answer” → gpt-5.1 / gpt-5-mini / Claude Sonnet 4.6
  • “It should be as fast and inexpensive as possible” → gpt-5-nano / gpt-4o-mini / Gemini 3.1 Flash-Lite / Claude Haiku 4.5
  • “I want to program / write code” → Claude Sonnet 4.6 / gpt-5.4
  • “It’s complicated or extremely important” → gpt-5.4 / gpt-5.1
  • “Data protection (EU/Germany) is mandatory” → STACKIT models (e.g., Llama 3.3 70B Instruct)
  • “I don’t know which model fits” → auto (dynamic routing)
  • “I need images” → GPT Image 1.5 / Gemini Image Tools / Nano Banana