Load Balancer
Distribute requests across multiple inferencers.
A Load Balancer distributes incoming requests across multiple inferencers according to the balancing strategy you define.
- Distribute requests across multiple inferencers using custom Weights.
- Balance traffic with Random or Round Robin Algorithm.
- Use different inferencers together within a single request path.
How requests are distributed
The Algorithm determines how requests are distributed between inferencers, while Weights determine how much traffic each inferencer receives.
- Random makes a new random selection for every request, with Weights determining the probability of each inferencer being selected.
- Round Robin follows a repeating rotation between inferencers, with Weights determining how often each inferencer appears in the rotation.
Add your Targets
Targets are the inferencers this Load Balancer distributes requests across. Each target is an inferencer with its own Weight.
- Weights determine each target’s share of the total traffic relative to the other targets.
- A target can be any inferencer, including a Model, Load Balancer, Failover, Race, Router, or Gateway.
Expand what your Load Balancer can do
You can extend the Load Balancer with functionality for validation, cost management, conversation context, data cleanup, and debugging. Your application keeps calling the same endpoint, while each addition handles its specific responsibility on every request.
- A Validator makes sure every answer arrives in the format you expect, such as clean JSON.
- A Budget caps what the Load Balancer can spend per run, chat, day, week, or month.
- A Memory lets conversations remember earlier context, while a Wiper clears stored data according to your policy.
- A Log Group keeps a record of requests for debugging and review.
Select the functionality you need, and click Apply.
What a Load Balancer changes for your application
A Load Balancer gives your application one endpoint for multiple inferencers, while letting you decide how requests are distributed between them.
Distribute traffic across inferencers
Every request is assigned according to the Algorithm and Weight configured for the Load Balancer, allowing you to distribute traffic across multiple inferencers.
Rebalance without code changes
Add or remove a target, or change its Weight in the console. Subsequent requests follow the new distribution without changing your application code.
Any inferencer can be a target
Start with two Models and later replace one with a Failover or Race. The Load Balancer can distribute requests across any inferencers you choose.
A Load Balancer can also stand inside any other inferencer. Use it as a step in a Failover, a contender in a Race, or a route in a Router.
Ready to build your AI execution flow?
Configure your first inferencer with help from our team.