Sentinel Sidecar: a learning security layer for any HTTP service
A security sidecar you put in front of any HTTP service. OWASP rules and small ML models decide each request in about a millisecond, and an out-of-band LLM analyst turns what it learns into short-lived blocks.
- state
- deploying
- stage
- growing
- status
- open source
- started
- 2026.09
- source
- github ↗
Sentinel Sidecar is one extra container that sits between your app and the internet and decides, for every request, whether to let it through. The app itself needs no code changes: you stop publishing its port, and the sidecar becomes the only way in.
Each request gets one of four verdicts:
- allow: forwarded to the app.
- log: forwarded, but flagged as suspicious.
- challenge: a JavaScript proof-of-work page (HTTP 429). Solving it sets a signed clearance cookie.
- block: rejected with a 403.
A plain WAF matches known-bad signatures and stops there. This one adds behaviour and bot scoring, an ML payload classifier, prompt-injection defence for LLM endpoints, masking of leaked errors in responses, and an analyst that studies traffic over time and feeds what it finds back into the inline decision.
Source code: github.com/sumit-kumar-03/sentinel-sidecar
How it works
There are two lanes.
The inline lane decides every request before the app sees it. Envoy normalizes the request and runs
OWASP CRS 4.14 (through the Coraza WASM filter), per-IP rate limits and IP lists. It then asks the engine
over Envoy's ext_authz hook. The engine extracts features once, runs four small models and returns a verdict.
The out-of-band lane never sits on the request path. Every decision is streamed to Redis Streams, where an analyst looks for campaigns across many clients and asks a local LLM about the gray-zone requests. Its findings come back as time-limited denylist entries that the engine enforces on the very next request.
| Model | Looks for | Lane | How |
|---|---|---|---|
| L0 | Known-bad signatures | inline | OWASP CRS, rate limits, IP allow and deny lists |
| M1 | Payload attacks | inline | Character-level CNN, INT8 ONNX, about 0.3 ms: SQLi, XSS, traversal, command and template injection |
| M2 | Unusual behaviour | inline | Per-client request rate, distinct-path fan-out, parameter novelty, kept in a Redis window |
| M3 | Bots and scanners | inline | Scanner user-agents, missing browser-shaped headers |
| M4 | Prompt injection | inline | A second char-CNN, active only on routes marked as LLM endpoints |
| M5 | Gray-zone triage | out of band | A local instruct model through Ollama, with strictly validated JSON output |
| C1 | Campaigns | out of band | Distributed credential stuffing and slow scans across many IPs |
The inline models run on CPU. If a model file is missing, a regex fallback keeps the engine working.
Design choices worth knowing
- Scores combine with a weighted noisy-OR, not an average. Each model can only raise the risk. With an average, adding a model that has no opinion would dilute one that is confident, and that was a real regression during the build.
- Hard signals override the blend. A confident payload match, a known scanner, an extreme scan rate, a prompt injection or an active denylist entry decides the verdict on its own.
- Three modes for a safe rollout.
monitorscores and logs what it would have done,learnalso records per-route parameter baselines, and onlyenforceactually blocks. - Fail open or fail closed, per route. If the engine is unreachable, a login route can refuse traffic while a public page keeps serving.
- The client IP comes from Envoy, not from the request. Trusting
X-Forwarded-Forwould let an attacker rotate fake IPs to dodge the per-client limits, so the engine reads the address Envoy has already sanitized. - CRS runs in score mode by default. In block mode at paranoia level 2 it flagged about a quarter of normal API requests, so CRS contributes a signal and the engine makes the call.
- An invalid config never starts. A one-shot container validates
sentinel.yamland renders the Envoy and engine configs. If validation fails, the gateway doesn't start at all.
The feedback loop
A confident M5 verdict or a detected slow scan becomes a denylist entry in Redis with an escalating but capped TTL.
Guardrails keep it from going wrong: allowlisted clients are never denied, distributed campaigns only produce
proposals (never an automatic mass block), and every action is written to an audit stream. An internal admin API
lists the actions, active blocks and proposals, and lifts a block with one DELETE.
Run the demo
You need Docker with Compose v2.20 or newer (for include:). The demo puts OWASP Juice Shop behind the sidecar:
git clone https://github.com/sumit-kumar-03/sentinel-sidecar.git
cd sentinel-sidecar
make demo-up # builds the images and starts Juice Shop behind the sidecar on :8080
make smoke # 30 end-to-end checks against the live stack
The demo starts in monitor mode, so attacks are scored and logged but not blocked:
curl -s -o /dev/null -w '%{http_code}\n' 'http://127.0.0.1:8080/rest/products/search?q=apple'
# 200
# see the latest decision the engine recorded
docker compose -f deploy/compose/demo/compose.yaml exec sentinel-redis \
redis-cli XREVRANGE sentinel:events + - COUNT 1
Set mode: enforce in deploy/compose/demo/sentinel.yaml to watch it bite: a SQL injection or a scanner
user-agent gets a 403, and a burst of distinct paths from one client gets the 429 proof-of-work page. make demo-down
removes everything.
Protect your own app
-
Build the images once with
make images. -
Write a
sentinel.yamlnext to your compose file. Onlyupstream.urlis required, and the repo'ssentinel.example.yamldocuments every setting. -
Include the sidecar's compose fragment and remove your app's
ports:, so traffic can only enter through the sidecar:include: - path: /path/to/sentinel-sidecar/deploy/compose/sentinel.compose.yaml project_directory: . services: app: image: my-app # no ports here -
Start in
monitor, watch the events and the Prometheus metrics for a while, then switch toenforce.
The default lite profile is CPU only. The full profile adds Ollama for the M5 analyst. For Kubernetes, the repo has an example that runs the same containers as sidecars in the app's pod.
Results
make gate starts a throwaway stack in enforce mode, replays a labelled corpus through it and fails if any target
is missed:
| Check | Target | Result |
|---|---|---|
| Attack detection | 95% or more | 96.4% (54 of 56 payloads, 5 attack families) |
| False positives | 0.1% or less | 0% (0 of 2,000 benign inputs) |
| Added latency | under 25 ms at p95 | about 1.4 ms at p95 |
| ffuf directory brute force | blocked | 60 of 60 requests blocked |
| sqlmap | no injection found | neutralized: 403s, parameter reported not injectable |
The corpus is small and hand-picked, so treat these numbers as a regression gate rather than a benchmark. There are also 101 unit tests and a 30-check smoke suite.
Limits and next steps
- The models are trained on synthetic data. Detection drops to about 80% on deliberately evasive payloads. Training on real corpora such as SecLists, CSIC 2010 and live traffic is the next step.
- CRS and the engine don't share scores yet. Feeding the CRS anomaly score into the engine's blend should raise detection without CRS's false positives.
- Still to build: per-account rate limits, bounded threshold tuning from the analyst, a gRPC
ext_authztransport, and masking only the leaked part of a response instead of the whole body.
Connected
shares a topic with this projectStreamlit Portfolio Template
A portfolio and resume website written entirely in Python. Edit one data file, run one command, and you have a dark-themed, animated personal site with a working contact form.
Botly: a local RAG chatbot
A private chatbot that answers from your own PDFs. It runs entirely on your machine with Ollama, LangChain, FAISS and Streamlit, packaged in one Docker image.
AppSec scanner suite: six scanners, one interface
SAST, DAST, SCA, SBOM, CSPM and secret detection, each wrapping a proven open-source scanner behind the same command line, the same Docker packaging and the same JSON output.
CVE Trove: a vulnerability intelligence pipeline
Pulls 15 public vulnerability feeds on a schedule, merges them into one record per CVE, and streams the result into MongoDB through Celery workers.