Race

Get the fastest successful response.

A Race sends the same request to multiple inferencers at the same time and returns the first successful response.

  • Run multiple inferencers simultaneously.
  • Use different inferencers as competing paths for the same request.
  • Return the first successful response.
ApplicationRequestFastest replyRaceCalls all at onceModel 1Model 2Model 3

General
Targets
Common
Tags
More
*Name
Fastest of Two
Description
Please enter descriptionBoth models run, the faster answer wins.
Race applied
Apply

Define your Race

A name and description identify the Race and make it easier to understand when managing multiple execution paths. The actual contenders are defined through its Targets.


Add your Targets

Every target receives the request at the same time, and the fastest response goes back to your application.

  • Add as many targets as you need when speed matters most.
  • A target can be any inferencer, including a Model, Load Balancer, Failover, Router, or another Race.
  • No matter how many targets participate in the Race, it counts as one API call.
General
Targets
Common
Tags
More
*Targets
Type
Please select oneMODEL▼

MODEL
LOAD_BALANCER
FAILOVER
RACE
ROUTER
GATEWAY
Model 1
Please select oneFast Chat Model▼

Fast Chat Model
Fast Drafting Model
Type
Please select oneMODEL▼
Model 2
Please select oneFast Drafting Model▼
Race applied
Apply

General
Targets
Common
Tags
More
Validator
Please select oneJSON Response Validator▼

JSON Response Validator
YAML Config Validator
Budget
Please select oneMonthly Cost Cap▼

Monthly Cost Cap
Daily Spend Guard
Memory
Please select oneSupport Chat Memory▼

Support Chat Memory
Wiper
Please select oneThirty Day Cleanup▼

Thirty Day Cleanup
Log Group
Please select oneProduction Logs▼

Production Logs
Applied
Apply

Expand what your Race can do

You can extend the Race with functionality for validation, cost management, conversation context, data cleanup, and debugging. Your application keeps calling the same endpoint, while each addition handles its specific responsibility on every request.

  • A Validator makes sure every answer arrives in the format you expect, such as clean JSON.
  • A Budget caps what the Race can spend per run, chat, day, week, or month.
  • A Memory lets conversations remember earlier context, while a Wiper clears stored data according to your policy.
  • A Log Group keeps a record of requests for debugging and review.

Select the functionality you need, and click Apply.


What a Race changes for your application

Get the fastest response from multiple inferencers running at the same time, with one API call and all results available for comparison.

The fastest response wins

Every contender receives the request at the same time, and the fastest successful response is returned to your application.

Multiple inferencers, one API call

The Race can run as many targets as you need, while your application sends only one API call to the Race.

Compare every result

The Race keeps the results from all contenders, so you can compare their responses and see how they performed against the winner.

A Race can also stand inside any other inferencer. Use it as a target of a Load Balancer, a step in a Failover, or a route in a Router.



Ready to build your AI execution flow?

Configure your first inferencer with help from our team.