Load Balancer Console Guide

This page provides a guide to the Load Balancer entity in the Namirasoft Inference Console. It defines the concepts and configuration fields used when creating and managing Load Balancers. Use this guide to understand the purpose and behavior of each setting available during Load Balancer configuration.

What Is a Load Balancer?

A Load Balancer in Namirasoft Inference distributes requests across several targets, so that no single target carries all the load. Each target is an inferencer, which can be a Model or any other inferencer, and each is given a weight that sets how much traffic it receives.

Any application can call a Load Balancer through the Namirasoft Inference API, in the same way as a Model. When a request arrives, the Load Balancer selects one of its targets using the configured algorithm and the target weights. Namirasoft applications such as Namirasoft Expert and Namirasoft Job Arranger also use it.

The Challenge a Load Balancer Solves

Sending every request to a single model concentrates load in one place, can reach a provider’s usage limits, and leaves no room to balance cost and speed across models.

Common challenges include:

  • Concentrated load: A single model handles all traffic, with no way to spread requests across several models. Busy periods hit one place, and everything behind it slows together.
  • Provider limits: High volume on one model can reach the usage limits of a single provider. When a limit is reached, every request feels it at once.
  • Cost and speed trade-offs: Without weighting, there is no simple way to send more traffic to a cheaper or faster model. The mix of cost and speed you want becomes a decision you cannot express.

How Namirasoft Inference Solves the Problem

A Load Balancer lets you group several targets and share requests across them. You give each target a weight and choose an algorithm that decides how the next request is routed. You can send more traffic to a faster or cheaper model, or spread traffic evenly to stay within each provider’s limits, all without changing the applications that use the Load Balancer.

Overview of Load Balancer Fields and Options

Below is a detailed explanation of the fields available when creating or managing a Load Balancer. Understanding these fields helps ensure your Load Balancer is configured correctly for your requirements.

  • ID (String): This is a unique identifier automatically assigned to the Load Balancer when it is created. The system uses it to track, reference, and manage this specific Load Balancer. This value is auto-generated and cannot be modified.
  • User ID (Namirasoft Account’s ID): This is the unique identifier of the Namirasoft Account user who owns this Load Balancer. It is used internally for permission control, audit logging, and access management.
  • Workspace ID (Namirasoft Workspace’s ID): This is the identifier of the workspace this Load Balancer belongs to, as defined in Namirasoft Workspace. A workspace is a shared organizational space where teams group their configurations, projects, and members.
  • Name (String): This is a label used to identify this Load Balancer in the console. A good name clearly describes its purpose, for example: “Primary Balancer” or “Cost Optimized Pool”.
  • Algorithm (Enum): This determines how requests are distributed between targets, while the Weights determine how much traffic each target receives.
    • Random: This makes a new random selection for every request, with the Weights determining the probability of each target being selected.
    • Round Robin: This follows a repeating rotation between targets, with the Weights determining how often each target appears in the rotation.
  • Description (String): This is a note that describes the Load Balancer and its purpose. It is for your reference and does not affect routing.
  • Targets (List): These are the inferencers this Load Balancer distributes requests across. Add one or more targets, each with the following:
    • Type (Enum): This is the kind of target, one of Model, Load Balancer, Failover, Race, Router, or Gateway. Because a target can itself be another execution flow, Load Balancers can be combined with other inferencers.
    • Target: This is the specific inferencer, of the chosen Type, that receives requests.
    • Weight (Number): This sets the relative share of traffic this target receives. A higher weight means more requests are routed to it.
  • Budget (Budget’s ID): This determines whether spending through the Load Balancer is subject to a defined limit. A Budget can set spending ceilings per run, chat, day, week, or month. The same Budget can be used by several inferencers.
  • Memory (Memory’s ID): This determines whether the Load Balancer keeps conversation context across requests. When a Memory is connected, relevant conversation history is stored and provided as context when the Load Balancer handles later requests.
  • Wiper (Wiper’s ID): This determines whether stored data associated with the Load Balancer is cleaned up according to a defined policy. A Wiper can remove data based on time or count limits.
  • Log Group (Log Group’s ID): This determines where the Load Balancer’s execution logs are recorded. When a Log Group, defined in Namirasoft Log, is connected, request activity and errors are recorded there for debugging and review.
  • Created At (DateTime): This is the date and time when this Load Balancer was created. This value is automatically generated and cannot be modified.
  • Updated At (DateTime): This is the date and time when this Load Balancer configuration was last modified. This value is updated automatically whenever any field is changed.


Ready to Configure Your Load Balancer or Want to Explore the Console Guide?