> ## Documentation Index
> Fetch the complete documentation index at: https://docs.kaireonai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Logistic Regression

> Weighted sum of features → sigmoid. Calibrated probabilities with linear feature interactions. Fast, well-understood, L2-regularizable.

`modelType: "logistic_regression"` — a single-layer linear model: dot-product the customer's feature vector with a learned weight vector, add a bias, push through a sigmoid. Probably the most-deployed classifier in production decisioning systems for a reason: cheap to train, cheap to score, easy to defend.

## When to use

* **You have ≥ 1k labeled outcomes and numeric features** — the linear weighted-sum structure benefits from numeric inputs (categoricals need one-hot).
* **You need calibrated probabilities for budget pacing or expected-value calculations** — the sigmoid output is calibrated within the linear region.
* **You're comparing against a Bayesian baseline** — logistic and Bayesian are the two "first-real-model" picks. Train both, A/B test them via `shadowModelKeys[]`.

**Skip it when** features have meaningful interactions (e.g. "income matters more for high-credit-score customers") — linear regression can't capture interaction terms without feature engineering. Use `gradient_boosted` instead.

## The math

```
z      = bias + Σ_i (weights[xᵢ] × xᵢ)
score  = sigmoid(z) = 1 / (1 + e^(-z))
```

Training minimizes binary cross-entropy with optional L2 regularization via full-batch gradient descent, running for `maxIterations` passes over the data (default 100).

## Fixture config

```json theme={null}
{
  "modelType": "logistic_regression",
  "predictors": [
    { "field": "credit_score", "selected": true, "importance": 0.40 },
    { "field": "income",       "selected": true, "importance": 0.25 },
    { "field": "age",          "selected": true, "importance": 0.10 },
    { "field": "segment",      "selected": true, "importance": 0.25 }
  ],
  "modelState": {
    "weights": {
      "credit_score": 0.005,
      "income":       0.00002,
      "age":          0.01,
      "segment":      0.5
    },
    "bias": -4.0
  }
}
```

The algorithm-coverage proof verifies this scores `0.930` for the standard test customer. Highest contribution: `credit_score × 760 × 0.005 = 3.8` (raw); next `income × 95000 × 0.00002 = 1.9`.

## Training

`POST /api/v1/algorithm-models/{id}/train` runs batch gradient descent over the observed interactions, minimizing binary cross-entropy to converge the weights. The training routine's hyperparameters live on `model.config`:

Before extracting samples, training merges **schema-table customer enrichment** (`ds_*` tables) into each interaction's feature bag — the same bulk enrichment load the Bayesian, gradient-boosted, and online-learner trainers use. A logistic model whose predictors reference schema columns (e.g. `credit_score` from a customer schema) therefore trains on those features even when the interaction rows' `context` doesn't carry them. On key collision, the interaction row's own `context` wins over the enriched attributes. (Previously, schema-column predictors produced zero usable samples and training silently fell back to the metrics-only path while still reporting a successful train.)

```json theme={null}
{ "learningRate": 0.01, "maxIterations": 100, "regularization": "l2", "regularizationStrength": 1.0 }
```

| Config key               | Default | Meaning                                             |
| ------------------------ | ------- | --------------------------------------------------- |
| `learningRate`           | `0.01`  | Step size per gradient-descent update.              |
| `maxIterations`          | `100`   | Number of full-batch passes over the training data. |
| `regularization`         | `"l2"`  | Penalty family: `none`, `l1`, or `l2`.              |
| `regularizationStrength` | `1.0`   | Penalty magnitude (λ).                              |

<Note>
  **Real fit vs. metrics-only fallback.** The engine fits genuine weights (consuming `learningRate` and `maxIterations`) only when it has **≥ 20 usable labeled rows** with **both classes present** (at least one positive *and* one negative outcome). Below that — too little signal to fit a stable model — training falls back to a **metrics-only** pass: it records evaluation metrics but leaves the weight vector unchanged rather than producing a degenerate all-one-class model. Collect more labeled outcomes across both classes to cross the threshold.
</Note>

<Note>
  **How the L2 penalty is wired.** The batch fitter reads `learningRate` and `maxIterations` from `config` directly, and resolves the L2 penalty strength (λ) from the enum-style config:

  * `regularization: "l2"` → λ = `regularizationStrength` (default `1.0`)
  * `regularization: "none"` → λ = `0` (no penalty)
  * `regularization: "l1"` → **not implemented** in the batch fitter, which only applies an L2 penalty. The engine logs a warning and applies λ = `0` rather than mislabeling the result as L1-regularized. Use `"l2"` if you want a penalty.
  * A **numeric** `regularization` value is still honored directly as λ (legacy back-compat), and takes precedence over `regularizationStrength`.

  A training run left on the default config (`regularization: "l2"`, `regularizationStrength: 1.0`) is therefore L2-regularized with λ = 1.0. Raise `regularizationStrength` to penalize large weights harder; set `regularization: "none"` to fit without a penalty.
</Note>

Categorical features need explicit one-hot expansion in the training data — the engine binarizes `segment="Gold"` to `1` if present, `0` if absent. Multi-valued categoricals (`segment` ∈ {Bronze, Silver, Gold, Platinum}) need 4 binary features.

## Score interpretation

* `score` ∈ `[0, 1]` — calibrated probability.
* `explanations[]` — per-feature contribution `weight × value`, sorted by absolute magnitude. Positive contributions push toward responding.

## Pitfalls

* **Categoricals treated as scalars** — `segment = 3` for Gold is nonsense (no ordinal relationship). Always one-hot expand.
* **Unscaled features** — `credit_score` (300–850) and `income` (0–500000) on the same model dominate `age` (18–80). Standardize to z-scores or min-max normalize before training, otherwise weights for small-magnitude features get pushed to zero by L2.
* **Multicollinearity** — heavily correlated features split the credit; explanations become misleading. Drop one of each correlated pair.
* **Missing intercept** — leaving `bias` at 0 forces every score through the origin. Always include the bias term.
* **Class imbalance** — if positive rate is 1% and the loss is unweighted, the model learns to always predict "negative". Use class weighting or downsample negatives in training.

## Lifecycle & cadence

Logistic regression is **offline-retrain only**. The weight vector and AUC are recomputed by `executeRetrain` against accumulated `interaction_history`. With `autoLearn: false` (the default), the model stays frozen — weights from the last manual training run remain in `modelState.weights` forever and `metricsHistory` doesn't grow.

To enable: `PUT { "autoLearn": true, "learnMode": "scheduled", "learnSchedule": "24h" }`. Nightly is the standard cadence; drop to `"6h"` if your underlying distribution shifts within a day, raise to `"7d"` for very stable workloads. Each retrain writes a fresh entry to `metricsHistory[]` and updates `lastTrainedAt` + `lastLearnedAt`. See [Learning cadence](/ai-ml/learning-cadence) and [Model lifecycle](/ai-ml/model-lifecycle).

## Cross-reference

* [Algorithm Selection Guide](/decisioning/algorithm-selection-guide).
* [Bayesian](/ai-ml/algorithms/bayesian) — natural baseline comparison.
* [Gradient Boosted Trees](/ai-ml/algorithms/gradient-boosted) — pick this instead when interactions matter.
