This is the first real build post in the series, so it starts with the architecture design. Working with an LLM through this process felt a lot like having an extra pair of hands on the team — but only as good as the context and constraints I actually gave it.

I went in half-expecting the AI to just write the whole thing. It didn’t work that way. Generated code is only as good as the engineering judgment steering it — without solid fundamentals and domain knowledge to check the output against, you end up shipping whatever the model happened to produce.

Series Name: Bastion-RAG – Project Security RAG

  • [Bastion-RAG] Project Security RAG
  • [Bastion-RAG 0] Get help from AI (Architecture Design) – Here!
  • [Bastion-RAG 1 – Sentinel]
    • Prompt Injection Defense
    • Metadata Filtering
  • [Bastion-RAG 2 – Vault]
    • Multi-tenancy
    • Deterministic De-identification
  • [Bastion-RAG 3 – Navigator]
    • Hybrid Reranking
    • Logical Partitioning
  • [Bastion-RAG 4 – Archor]
    • Embedding Noise Injection
    • Embedding Model Bias Verification
  • [Bastion-RAG 5 – Tracker]
    • Data Lineage Tracking
    • Honey-token Injection
  • [Bastion-RAG Demo]


The Bastion-RAG project was initiated to implement security governance layers—including prompt injection defense, deterministic tokenization, and differential noise injection—into a production-ready system. Initial development assumed that integrating these security features would be straightforward given the capabilities of modern AI coding assistants.

While initial code generation was completed quickly, architectural implementation introduced complex design variables. Establishing an architecture that satisfies enterprise compliance mandates (such as PIPA and GDPR) while minimizing runtime latency on live data pipelines posed a significant challenge. Appending security control layers naively increased system latency beyond acceptable metrics. Resolving these performance bottlenecks required three structural iterations of the core framework and extensive design validation with the AI assistant.

This post serves as an engineering retrospective detailing how the pipeline’s structure was refined through automated validation and meticulous execution overhead measurement to construct a secure RAG governance framework.

Architecture Design

2. Evolutionary Analysis: The Dramatic Structural Journey from v1 to v3

Reviewing the architectural design prompt records exchanged with the AI assistant reveals a continuous cycle of technical friction, rapid prototyping, and a stark realization of my own initial research gaps.

2.1 [Version 1.0] The Swamp of Functional Fragmentation and Asymmetry (Failure)

  • Initial Architectural Profile
    • The initial iteration implemented a unidirectional defense architecture focused primarily on the input pipeline. Each security component, including Sentinel-IN and Vault-IN, operated as a decoupled microservice or independent container. The Honey-Token intrusion detection system was also designed as an isolated function confined within the Tracker module.
  • Design Intent and Architecture Selection
    • The initial objective was to decouple the security modules to ensure system extensibility. The architecture featured an inbound gateway written in Go, with the embedding generation and reranking components isolated in a separate Python model server accessed via HTTP calls. This approach adopted standard microservice principles to establish clear deployment boundaries across domains.
  • Performance and Security Evaluation
    • Integration benchmarks revealed significant performance overhead. Each retrieval request triggered multiple external microservice calls, accumulating HTTP network hop delays and serialization/deserialization overhead. As a result, p95 latency increased beyond 300ms, failing to meet real-time processing requirements.
    • Additionally, benchmarking exposed a structural asymmetry in the pipeline. The exclusive focus on inbound validation resulted in the omission of an output verification phase. This created a security risk where the LLM could include raw personal data or unauthorized enterprise secrets in its final response, bypassing storage-level isolation mechanisms.

2.2 [Version 2.0] Discovery of Symmetrical Integration and Cross-Cutting Dynamics (Transition)

  • Revised Architectural Profile
    • The second iteration consolidated the separate input (Phase 1) and output (Phase 2) modules into a bidirectional symmetric architecture within a single Go service container. Features such as Honey-Tokens, Multi-Tenancy, and Data Lineage were refactored from isolated modules into cross-cutting system-wide coordinators.
  • Language Ecosystem Limitations
    • Consolidating the pipeline into a single Go process eliminated network hop latency, but introduced limitations in the machine learning operation layer. The Go ecosystem lacks native primitives for operations like Word Embedding Association Tests (WEAT) or Laplacian noise distributions, requiring custom low-level implementation. Strict single-language consolidation created a maintenance burden and limited architectural flexibility. The system required a hybrid interface contract capable of maintaining performance isolation while leveraging the Python machine learning ecosystem.
  • Limitations of the Transition
    • While this single-language consolidation resolved the network latency bottleneck, it introduced maintenance challenges due to machine learning serving overhead and complex CGO bindings.

2.3 [Version 3.0] Completion of the High-Performance Polyglot Wire Contract (Current)

  • Final Architectural Profile
    • The final iteration retains the symmetrical dual-phase pipeline and decoupled module design established in v2.0, implementing an optimized hybrid polyglot structure. High-speed text pattern matching and cryptographic token mapping are handled by Go, while vector embeddings and matrix numerical operations are managed entirely by Python.
  • Inter-Process Communication Optimization
    • To leverage the Python machine learning ecosystem without introducing the network communication latency observed in v1.0, the Navigator (Search) and Anchor (Security) modules were redesigned as self-contained Python processes. Hosting the sentence-transformers and CrossEncoder models directly within process memory allows inference to be executed in-process, eliminating network hops.
    • To bridge the multi-language boundary, the system utilizes a gRPC infrastructure. The standard binary protobuf layer is replaced with a custom JSON Codec contract on the wire, balancing structural flexibility with strict type safety.

3. [Bastion-RAG 0] Virtual Emulation Simulation for Architecture Design

Through three architectural iterations, implementation demonstrated the necessity of validating configuration compliance and exception handling paths prior to codebase development. To integrate this approach into the development lifecycle, Bastion-RAG 0 was established as a proactive architectural auditing layer at the framework’s entry point.

When enterprise security policies and compliance constraints are defined, Bastion-RAG 0 emulates the pipeline’s event stream over a NATS event bus topology. This emulation runs without loading machine learning models or initializing concrete services.

Once a proposed pipeline configuration is provided, the Bastion-RAG 0 audit engine dynamically generates a real-time Virtual Ingestion Trace Log to analyze data lineage flows and identify potential architectural bottlenecks:

[Bastion-RAG 0: Virtual Emulator Ingestion Trace Log]

 - [emulator/Sentinel-IN]  INFO: Prompt input validated. Status: PASSED. Injection score: 0.05[cite: 7]
 - [emulator/Vault-Phase1] INFO: Multi-strategy anonymization executed.[cite: 6]
                               - Input matching: "Hong Gildong" -> KR_NAME_8f3d2a (PERSON)[cite: 7]
                               - Input matching: "[email protected]" -> EMAIL_c3a91f (EMAIL)[cite: 7]
 - [emulator/Navigator]    INFO: Executing structural pre-filtering isolation.
                               - Injected filters: tenant_id=acme, collections=[customer_docs][cite: 7]
 - [emulator/Anchor-IN]    INFO: Differential noise injection executed. Sigma applied: 0.01
 - [emulator/Vault-Phase2] INFO: Evaluating selective detokenization via OPA policy rules.
 - [emulator/Sentinel-OUT] INFO: Grounding and hallucination checks completed. Status: PASSED.

4. Architectural Principles and Hard Constraints for Total Isolation

The constraints below are written the way an actual spec should read: plainly, without the dramatic framing an AI assistant tends to reach for by default.

They exist to be checked against, not read as prose.

The architectural constraints of the Bastion-RAG framework are defined as follows:

Core Functional Autonomy

Each module must be designed as a self-contained, autonomous unit that delivers security value independently when integrated with an LLM. To ensure graceful degradation, the failure of peripheral modules must not disrupt or cause cascading errors in the primary data path.

Prohibition of Direct Coupling

Modules must not instantiate or execute direct API calls to other modules. For example, the Navigator search layer does not maintain a reference to the Vault. It operates on a zero-trust data contract, where required user permission tokens are retrieved by the upstream orchestrator and included directly in the request payload.

Non-Invasive Observability Architecture

The Tracker module, which aggregates system audit records and maps data lineage paths, must not introduce synchronous blocking or latency to the primary data path. It functions as a non-invasive observer, consuming asynchronous, fire-and-forget JSON event streams transmitted over a decoupled NATS message bus.

5. Conclusion: The AI Assistant Is an Incredible Debating Partner, Not a Blindly Trusted Tool

Getting here took real iteration — the textbook microservice pattern simply couldn’t hit the latency numbers this needed. What worked was using the AI assistant to hold the line on explicit latency targets and language boundaries, which is how the pipeline ended up at the polyglot v3 design.

Prototyping took about four weeks. One early version shipped with a dependency flaw that forced a rollback and a two-week refactor. That detour was its own lesson in the practical limits of working with an LLM — token limits and context window degradation are easy to underestimate until you hit them.

It reminded me of earlier refactoring work — shrinking a legacy codebase to a fifth of its size, and turning a month-long manual deployment process into a three-hour automated pipeline. Same lesson both times: complex systems get tamed by structured re-engineering and real test coverage, not by any one clever trick.

Building Bastion-RAG was a good, concrete lesson in what working with an LLM actually looks like day to day: it doesn’t replace engineering fundamentals or architectural discipline, it just changes where you spend your attention.

By Mark

-_-