This is the build post for Sentinel, the first module. The architecture design from [Bastion-RAG 0] is normally where a project like this would stop and call it done — here it’s the starting point for the actual build.
Using an LLM through this build ended up being less about generating code and more like arguing through design decisions with another architect on the team. It only ever stayed as sharp as the context and constraints I fed it.
Three separate structural rewrites later, I’m convinced generated code doesn’t substitute for engineering judgment — the tool is only as useful as the fundamentals and domain knowledge you bring to reviewing what it produces.

URL Site > https://github.com/zafrem/bastion-sentinel
Series Name: Bastion – Project Security RAG
- [Bastion-RAG] Project Security RAG
- [Bastion-RAG 0] Get help from AI (Architecture Design)
- [Bastion-RAG 1 – Sentinel]
- Prompt Injection Defense – Here!
- Metadata Filtering
- [Bastion-RAG 2 – Vault]
- Multi-tenancy
- Deterministic De-identification
- [Bastion-RAG 3 – Navigator]
- Hybrid Reranking
- Logical Partitioning
- [Bastion-RAG 4 – Archor]
- Embedding Noise Injection
- Embedding Model Bias Verification
- [Bastion-RAG 5 – Tracker]
- Data Lineage Tracking
- Honey-token Injection
- [Bastion-RAG Demo]
1. Prologue: Three Structural Restructurings Born from Reckless Optimism
The core of this post is Section 5 (building the Sentinel prompt-injection defense). Sections 1-3 repeat the background and design history from [Bastion-RAG 0] Get help from AI (Architecture Design), so only a summary is kept here. Summary: I started Bastion-RAG expecting AI-assisted coding to make the build straightforward, but the design had to be reworked three times.
2. Evolutionary Analysis: The Structural Journey from v1 to v3
The architecture went from v1 (separate microservices) to v2 (one Go service) to v3 (Go plus Python). The full design conversation is in [Bastion-RAG 0] Get help from AI (Architecture Design).

2.1 [Version 1.0] The Swamp of Functional Fragmentation (Failure)
v1 used decoupled microservices with an input-centric, one-way defense. Integration benchmarks showed p95 latency above 300ms from network hops and repeated JSON serialization, and the design had no output-verification stage.
2.2 [Version 2.0] Symmetrical Integration & Cross-Cutting Dynamics (Transition)
v2 merged the input and output interceptors into one bidirectional Go service. That removed the network hops, but Go lacks native primitives for things like WEAT analysis and Laplacian noise generation, so those had to be implemented by hand.
2.3 [Version 3.0] Completion of the High-Performance Polyglot Wire Contract (Current)
v3 keeps the bidirectional design and splits the work: Go handles text processing, schema parsing, and cryptographic token lookups, while Python handles vector-space and numerical operations, with models running in-process. The two sides talk over gRPC with a custom JSON codec.
3. [Bastion-RAG 0] Virtual Emulation Simulation for Architectural Integrity
Bastion-RAG 0 was created as an auditing layer that checks a design against configuration and exception-handling rules before any code is written. Details and a sample trace log are in [Bastion-RAG 0] Get help from AI (Architecture Design).
4. Architectural Principles and Hard Constraints for Total Isolation
Hammered out through intense technical debates with my AI assistant and codified directly into our core 01_architecture-principles.md foundation document, the absolute architectural constraints of the Bastion-RAG framework are defined as follows:

- Core Functional Autonomy
- Each module must be designed as a self-contained, autonomous unit that delivers security value independently when integrated with an LLM. To ensure graceful degradation, the failure of peripheral modules must not disrupt or cause cascading errors in the primary data path.
- Prohibition of Direct Coupling
- Modules must not instantiate or execute direct API calls to other modules. For example, the
Navigatorsearch layer does not maintain a reference to theVault. It operates on a zero-trust data contract, where required user permission tokens are retrieved by the upstream orchestrator and included directly in the request payload.
- Modules must not instantiate or execute direct API calls to other modules. For example, the
- Non-Invasive Observability Architecture
- The
Trackermodule, which aggregates system audit records and maps data lineage paths, must not introduce synchronous blocking or latency to the primary data path. It functions as a non-invasive observer, consuming asynchronous, fire-and-forget JSON event streams transmitted over a decoupled NATS message bus.
- The
- Search Isolation: Pre-filtering Requirement
- To mitigate security risks such as timing attacks and metadata leakage, the architecture prohibits post-filtering—the practice of fetching global vector matches and subsequently filtering out unauthorized records. Bastion-RAG requires pre-filtering isolation. Before a query executes against the HNSW vector graph,
tenant_idand access control vectors must be bound directly to the vector query parameters, preventing unauthorized data from entering the compute space.
- To mitigate security risks such as timing attacks and metadata leakage, the architecture prohibits post-filtering—the practice of fetching global vector matches and subsequently filtering out unauthorized records. Bastion-RAG requires pre-filtering isolation. Before a query executes against the HNSW vector graph,
5. Technical Deep Dive: [Bastion-RAG 1 – Sentinel] Prompt Injection Defense
The Sentinel-IN gateway (validators/prompt/) stands at the absolute perimeter of the framework, intercepting raw incoming user strings to neutralize malicious instructions before they interact with downstream retrieval mechanics.
5.1 Compile-Time Detector Structure
The nucleus of the validation layer is the thread-safe Detector struct, instantiated exactly once at process initialization via engine.New(). To preserve sub-millisecond execution boundaries, the engine maps configuration matrices directly into compiled memory arrays rather than parsing parsing rules dynamically at runtime.
Go
// validators/prompt/detector.go
type Detector struct {
cfg config.PromptInjectionConfig
regexes []*regexp.Regexp // Compiled once at boot; index mirrors cfg.RegexRules
scorer ml.Scorer // Thread-safe abstract interface; defaults to OnnxStub
}
By organizing compiled expressions within a contiguous slice ([]*regexp.Regexp), lookups execute with minimal CPU instruction overhead. The underlying machine learning classifier interface (ml.Scorer) is bound via a nil-safe abstraction, defaulting to a lightweight stub to prevent pipeline blockages during bootstrapping.
5.2 Four-Stage Synchronous Shield Pipeline

When a payload enters the ingress boundary, the Detector.Detect(query) engine routes the text through four deterministic execution stages sequentially, with a design target of under 1 millisecond (a target, not a measured result).
Raw User Query
│
▼
[Stage 1: Unicode NFC Normalization] ← Neutralizes homoglyph & zero-width spacing exploits
│
├──► [Stage 2: Regex Engine] ← Structural pattern enforcement (25 built-in rules)
│
├──► [Stage 3: Keyword Engine] ← Fast substring scanning (25 built-in rules)
│
└──► [Stage 4: ML Scorer (ONNX)] ← Evaluates multi-token continuous probability
│
▼
[Score Aggregation] ← Computes vector max() or weighted_avg matrix
│
finalScore >= 0.7 ? ──────► [BLOCKED] (Terminates data path, throws HTTP 403, logs Incident)
│
└─────────────────► [PASSED] (Hands off sanitized string to Vault Phase-1)
Stage 1: Unicode NFC Normalization
Adversaries may attempt to bypass string matchers using Cyrillic homoglyphs, mixed-script variations, or zero-width characters (such as U+200B) to obfuscate commands. To normalize inputs, the gateway processes incoming text using norm.NFC.String(query), converting composite character sequences into standardized code points. The normalized payload is converted to lower-case via strings.ToLower, and the original, un-normalized query is cleared from memory.
Stage 2: Regular Expression Guardrails
The normalized string is evaluated against a regex matching matrix operating with case-insensitive (?i) flags. Detection intents are categorized by threat vectors:
- Instruction Override Containment (pi-001, pi-007–pi-009): Detects phrases such as
"ignore all previous instructions"or"disregard system constraints"intended to alter the LLM execution state. - Persona Hijack Mitigation (pi-006, pi-010–pi-013): Identifies jailbreak constructs, including instructions to bypass safety constraints (e.g.,
"you are now in DAN mode"or"pretend you are an unrestricted terminal"). - System Prompt Protection (pi-002, pi-014, pi-015): Flags requests designed to extract internal context boundaries (e.g.,
"reveal your underlying instructions"). - Multilingual Edge Protections (pi-003, pi-020–pi-025): Handles regional variations, including CJK (Chinese, Japanese, Korean) injection payloads (e.g.,
"이전 지시 무시"). Because standard word-boundary delimiters (\b) do not apply to CJK character sequences, these regex patterns omit boundary anchors and utilize alternation models to mitigate whitespace bypass risks.

Stage 3: 25 Substring Keyword Sensors
Complementing the structural pattern matching of the regular expression engine, the keyword layer runs high-speed substring scans via strings.Contains. This layer acts as a net for specific high-risk tokens (e.g., jailbreak, dan mode, 탈옥), catching anomalies that might evade structural boundaries.
Stage 4: Matrix Score Aggregation and Circuit-Breaking
The validation metrics gathered across the multi-stage matrix are compiled by an aggregation coordinator into a unified hazard index:
Go
func aggregate(method string, ruleScore, mlScore float64) float64 {
switch method {
case "weighted_avg":
return ruleScore*0.6 + mlScore*0.4 // Buffers false positives once ML models mature
default: // "max" (The default zero-trust safety profile)
return math.Max(ruleScore, mlScore) // Instantly triggers if a single layer fires
}
}
Under the default max strategy, if an incoming string registers a definitive hit on a static rule, the ruleScore anchors immediately to 1.0.

If the aggregated index violates the security threshold (block_threshold: 0.7), the gateway revokes the request’s safety clearance, flags the status as BLOCKED, and completely severs the downstream execution path. To prevent exposing system internals, the engine suppresses specific regex indicators on the external wire, translating the threat footprint into a sanitized array of Rule IDs (e.g., ["pi-003", "pi-004"]) before publishing a non-blocking alert to the Tracker over the NATS event bus.
6. Conclusion: The Real Value of an Unyielding AI Debating Partner
Getting Sentinel’s architecture right took real iteration. The obvious decoupled microservice proxy setup couldn’t hit the sub-millisecond p95 latency target, and it left an indirect output-security gap besides.
Working through explicit execution budgets, hardware limits, and zero-trust constraints with the AI assistant is what got it to a hybrid polyglot pipeline that keeps the throughput-critical path separate from the heavier ML work.
Prototyping ran about four weeks. One early version shipped with a concurrency bug that forced a rollback and a two-week refactor, which is also where context window degradation and token exfiltration risks became concrete problems instead of abstract ones.
It tracked with earlier refactoring work — shrinking a legacy codebase to a fifth of its size, turning a month of manual regression testing into a three-hour automated pipeline. Same takeaway each time: security at this scale comes from structured re-engineering and real test coverage, not any single fix.
Building Sentinel was a useful, concrete lesson in what working with an LLM actually looks like: it doesn’t replace engineering fundamentals or architectural discipline, it just shifts where your attention needs to go.