Skip to content
CortexDocs
Egress Gateway · Advanced routing

Advanced routing

Build requests, select by your own evaluations, and run multi-model deliberation in your application.

  • Request builderInspect the request before execution
  • Quality and costBalance the needs of your workload
  • Multi-model deliberationUnderstand each model's role

These recipes run in your application using ordinary Egress requests. Egress has no request-builder, quality-selection, or deliberation endpoint. The functions below build request bodies or call an already configured OpenAI client; their return values belong to your application.

Configure client using OpenAI Python setup. Choose organization-enabled models that accept Chat Completions. Each request still passes through the gateway's model and provider checks.

Request builder

Build independent request bodies before executing them. Preserve the prompt, conversation, and compatible request settings when changing the model.

Build request bodies

build-requests.py
from copy import deepcopy
 
 
def build_requests(request, models):
    if not models:
        raise ValueError("Choose at least one enabled model")
    return [dict(deepcopy(request), model=model) for model in models]

Call build_requests(request, models) with your native request and approved model IDs. No provider call occurs during construction. Inspect the bodies, then execute each with client.chat.completions.create(**body).

Read execution responses

The caller executes each body and receives an ordinary Chat Completions response. Keep each response's model, choices, and usage alongside its request and any execution error. The gateway does not return a combined requests or cost_latency envelope.

Quality and cost

Choose one request using your own measured quality and estimated cost. Each candidate row below is application data with model, quality, and estimated_cost_usd. Use enabled models, finite numeric values, one named evaluation and score scale, and cost estimates for the same workload. The public pricing catalog does not supply quality scores.

Select a request

choose-request.py
from copy import deepcopy
 
 
def choose_request(request, candidates, minimum_quality):
    eligible = [row for row in candidates
                if row["quality"] >= minimum_quality]
    if not eligible:
        raise ValueError("No candidate meets the quality threshold")
    chosen = min(eligible, key=lambda row: (
        row["estimated_cost_usd"], row["model"]
    ))
    return dict(deepcopy(request), model=chosen["model"])

This application rule chooses the lowest estimated cost among candidates meeting the threshold; model ID breaks equal-cost ties. It fails without sending a request when the candidate set is empty. It does not ask Auto Router to optimize a quality score.

Execute the selected request

Pass the returned body to client.chat.completions.create(**body). Record the evaluation source and workload estimate separately from the provider's actual usage. A cost estimate is not a billed result.

Multi-model deliberation

The application can call a small panel, ask an analyst to compare the replies, and request a final answer. This sequential example permits one to three panel models, disables SDK retries, and sets a per-call timeout. It accepts partial panel success, refuses an entirely failed panel, and propagates analyst or final-answer failures.

Run the model roles

deliberate.py
from copy import deepcopy
 
 
def deliberate(client, messages, panel_models, analyst_model, final_model):
    if not 1 <= len(panel_models) <= 3:
        raise ValueError("Choose one to three panel models")
    bounded = client.with_options(timeout=30.0, max_retries=0)
    panel = []
    for model in panel_models:
        try:
            response = bounded.chat.completions.create(
                model=model, messages=deepcopy(messages), max_tokens=512
            )
            text = response.choices[0].message.content
            if not text:
                raise ValueError("Panel returned no text")
            panel.append({"model": model, "response": response, "text": text})
        except Exception as error:
            panel.append({"model": model, "error": type(error).__name__})
 
    replies = [entry for entry in panel if "text" in entry]
    if not replies:
        raise RuntimeError("No panel response available")
    comparison = "Compare these candidate answers:\n" + "\n\n".join(
        entry["model"] + ": " + entry["text"] for entry in replies
    )
    analysis = bounded.chat.completions.create(
        model=analyst_model,
        messages=deepcopy(messages) + [{"role": "user", "content": comparison}],
        max_tokens=512,
    )
    analysis_text = analysis.choices[0].message.content
    if not analysis_text:
        raise RuntimeError("Analyst returned no text")
    final = bounded.chat.completions.create(
        model=final_model,
        messages=deepcopy(messages) + [{
            "role": "user",
            "content": "Use this comparison to answer the original task:\n" + analysis_text,
        }],
        max_tokens=512,
    )
    if not final.choices[0].message.content:
        raise RuntimeError("Final model returned no text")
    return {"panel": panel, "analyst": analysis, "final": final}

Call deliberate(client, messages, panel_models, analyst_model, final_model) with your configured client, original conversation, and compatible model IDs. The function invokes all roles itself; there is no special gateway deliberation payload.

Inspect the application result

Read result["final"].choices[0].message.content for the final text. Keep the panel's individual responses or errors and the analyst response for review. These dictionary keys are local application state, not an Egress response schema. Tool calls are not executed by this text-only recipe.

Up to five model calls compound latency and cost. Each successful response carries its own usage. Failed calls may incur charges that are not represented by returned usage. Add an overall deadline or parallel execution only when the application also defines cancellation and partial-result handling.

Choose the execution model

NeedApplication recipeModel calls
Inspect bodies before sendingRequest builderNone until the caller executes the bodies.
Select using measured quality and estimated costQuality and costOne after the caller executes the selected body.
Compare several replies and synthesizeMulti-model deliberationPanel calls, then analyst and final-answer calls.

Auto Router selects one target on a qualified request path. These recipes use explicit models and application-owned execution.

Continue with routing

  • Routing controls: distinguish managed settings from application policy and retries.
  • Auto Router: use managed single-target routing on qualified paths.
  • Egress Gateway: configure a client and make the first request.
Was this helpful?