When retrieved text becomes an instruction
A closer look at indirect prompt injection, the RAG trust boundary, and why retrieved content should never inherit authority.
Ideas investigated. Assumptions challenged. Lessons written down.
A closer look at indirect prompt injection, the RAG trust boundary, and why retrieved content should never inherit authority.
Follow one request from identity to resource access, and find the security decision that authentication alone cannot make.
A practical starting point for threat modeling: identify the assets, follow the data, and ask where trust changes.
Look past the tool description. What can the agent actually do, with whose identity, and under which constraints?
A small template for turning an interesting result into a useful experiment: setup, observation, explanation, and limits.
Understand direct and indirect prompt injection through a workplace example, practical defenses, and a safe trust-boundary exercise.
Follow private data through an AI assistant, distinguish training from retrieval, and practice checking who may see each record.
Learn how models, packages, adapters, and providers enter an AI system, and how to review their origin before trusting them.
Distinguish poisoned training data, tampered models, and corrupted retrieval sources, then practice tracing a false policy back to its origin.
Understand all ten OWASP LLM risks in the 2025 edition through plain-language examples, first defenses, and a small design exercise.
Follow a harmless-looking document into an AI assistant and learn where outside content can become an unauthorized instruction.
Keep information from tools separate from permission to act, with clear examples of provenance, narrow tools, and independent authorization.
Build a small, repeatable security evaluation that checks real outcomes, catches regressions, and explains what its scores mean.