all systems operationalIndia

Sentinel Sidecar: a learning security layer for any HTTP service

A security sidecar you put in front of any HTTP service. OWASP rules and small ML models decide each request in about a millisecond, and an out-of-band LLM analyst turns what it learns into short-lived blocks.

state
deploying
stage
growing
status
open source
started
2026.09
source
github ↗
topics

Sentinel Sidecar is one extra container that sits between your app and the internet and decides, for every request, whether to let it through. The app itself needs no code changes: you stop publishing its port, and the sidecar becomes the only way in.

Each request gets one of four verdicts:

  • allow: forwarded to the app.
  • log: forwarded, but flagged as suspicious.
  • challenge: a JavaScript proof-of-work page (HTTP 429). Solving it sets a signed clearance cookie.
  • block: rejected with a 403.

A plain WAF matches known-bad signatures and stops there. This one adds behaviour and bot scoring, an ML payload classifier, prompt-injection defence for LLM endpoints, masking of leaked errors in responses, and an analyst that studies traffic over time and feeds what it finds back into the inline decision.

Source code: github.com/sumit-kumar-03/sentinel-sidecar

How it works

Architecture: the inline lane and the out-of-band lane

There are two lanes.

The inline lane decides every request before the app sees it. Envoy normalizes the request and runs OWASP CRS 4.14 (through the Coraza WASM filter), per-IP rate limits and IP lists. It then asks the engine over Envoy's ext_authz hook. The engine extracts features once, runs four small models and returns a verdict.

The out-of-band lane never sits on the request path. Every decision is streamed to Redis Streams, where an analyst looks for campaigns across many clients and asks a local LLM about the gray-zone requests. Its findings come back as time-limited denylist entries that the engine enforces on the very next request.

ModelLooks forLaneHow
L0Known-bad signaturesinlineOWASP CRS, rate limits, IP allow and deny lists
M1Payload attacksinlineCharacter-level CNN, INT8 ONNX, about 0.3 ms: SQLi, XSS, traversal, command and template injection
M2Unusual behaviourinlinePer-client request rate, distinct-path fan-out, parameter novelty, kept in a Redis window
M3Bots and scannersinlineScanner user-agents, missing browser-shaped headers
M4Prompt injectioninlineA second char-CNN, active only on routes marked as LLM endpoints
M5Gray-zone triageout of bandA local instruct model through Ollama, with strictly validated JSON output
C1Campaignsout of bandDistributed credential stuffing and slow scans across many IPs

The inline models run on CPU. If a model file is missing, a regex fallback keeps the engine working.

Design choices worth knowing

  • Scores combine with a weighted noisy-OR, not an average. Each model can only raise the risk. With an average, adding a model that has no opinion would dilute one that is confident, and that was a real regression during the build.
  • Hard signals override the blend. A confident payload match, a known scanner, an extreme scan rate, a prompt injection or an active denylist entry decides the verdict on its own.
  • Three modes for a safe rollout. monitor scores and logs what it would have done, learn also records per-route parameter baselines, and only enforce actually blocks.
  • Fail open or fail closed, per route. If the engine is unreachable, a login route can refuse traffic while a public page keeps serving.
  • The client IP comes from Envoy, not from the request. Trusting X-Forwarded-For would let an attacker rotate fake IPs to dodge the per-client limits, so the engine reads the address Envoy has already sanitized.
  • CRS runs in score mode by default. In block mode at paranoia level 2 it flagged about a quarter of normal API requests, so CRS contributes a signal and the engine makes the call.
  • An invalid config never starts. A one-shot container validates sentinel.yaml and renders the Envoy and engine configs. If validation fails, the gateway doesn't start at all.

The feedback loop

Feedback loop: from finding to inline block

A confident M5 verdict or a detected slow scan becomes a denylist entry in Redis with an escalating but capped TTL. Guardrails keep it from going wrong: allowlisted clients are never denied, distributed campaigns only produce proposals (never an automatic mass block), and every action is written to an audit stream. An internal admin API lists the actions, active blocks and proposals, and lifts a block with one DELETE.

Run the demo

You need Docker with Compose v2.20 or newer (for include:). The demo puts OWASP Juice Shop behind the sidecar:

git clone https://github.com/sumit-kumar-03/sentinel-sidecar.git
cd sentinel-sidecar

make demo-up      # builds the images and starts Juice Shop behind the sidecar on :8080
make smoke        # 30 end-to-end checks against the live stack

The demo starts in monitor mode, so attacks are scored and logged but not blocked:

curl -s -o /dev/null -w '%{http_code}\n' 'http://127.0.0.1:8080/rest/products/search?q=apple'
# 200

# see the latest decision the engine recorded
docker compose -f deploy/compose/demo/compose.yaml exec sentinel-redis \
  redis-cli XREVRANGE sentinel:events + - COUNT 1

Set mode: enforce in deploy/compose/demo/sentinel.yaml to watch it bite: a SQL injection or a scanner user-agent gets a 403, and a burst of distinct paths from one client gets the 429 proof-of-work page. make demo-down removes everything.

Protect your own app

  1. Build the images once with make images.

  2. Write a sentinel.yaml next to your compose file. Only upstream.url is required, and the repo's sentinel.example.yaml documents every setting.

  3. Include the sidecar's compose fragment and remove your app's ports:, so traffic can only enter through the sidecar:

    include:
      - path: /path/to/sentinel-sidecar/deploy/compose/sentinel.compose.yaml
        project_directory: .
    
    services:
      app:
        image: my-app        # no ports here
    
  4. Start in monitor, watch the events and the Prometheus metrics for a while, then switch to enforce.

The default lite profile is CPU only. The full profile adds Ollama for the M5 analyst. For Kubernetes, the repo has an example that runs the same containers as sidecars in the app's pod.

Results

make gate starts a throwaway stack in enforce mode, replays a labelled corpus through it and fails if any target is missed:

CheckTargetResult
Attack detection95% or more96.4% (54 of 56 payloads, 5 attack families)
False positives0.1% or less0% (0 of 2,000 benign inputs)
Added latencyunder 25 ms at p95about 1.4 ms at p95
ffuf directory brute forceblocked60 of 60 requests blocked
sqlmapno injection foundneutralized: 403s, parameter reported not injectable

The corpus is small and hand-picked, so treat these numbers as a regression gate rather than a benchmark. There are also 101 unit tests and a 30-check smoke suite.

Limits and next steps

  • The models are trained on synthetic data. Detection drops to about 80% on deliberately evasive payloads. Training on real corpora such as SecLists, CSIC 2010 and live traffic is the next step.
  • CRS and the engine don't share scores yet. Feeding the CRS anomaly score into the engine's blend should raise detection without CRS's false positives.
  • Still to build: per-account rate limits, bounded threshold tuning from the analyst, a gRPC ext_authz transport, and masking only the leaked part of a response instead of the whole body.

Connected

shares a topic with this project