PROMPT INJECTION CHECK
What we detect & FAQ
A guide to the signals, file layers and limits of the current local scanner.
Prompt-injection signals
External content may include instructions aimed at an AI system rather than information for it to read. Rule-based checks look for supported instruction and role-manipulation signals; they do not provide a complete inventory of every possible wording.
Hidden and alternative text
Supported checks can surface concealed document text, very small or near-white text, hidden worksheets, comments, inline-hidden HTML, email preheaders and certain metadata or alternative textual channels. Coverage depends on the format and available extracted content.
Encoded and obfuscated content
Some supported transformed text, including bounded Base64 cases, can be interpreted for analysis. Invisible or directional characters may also be reported. A transformation by itself is not an instruction; review the decoded excerpt and source context.
Structured values
JSON, CSV, TSV and XML may carry instructions in fields, cells, attributes or nearby structured values. Supported XML CDATA and local fragments are inspected where available. Source locations help you return to the original data.
Supported formats and inspected layers
Examples below describe supported channels, not a guarantee that every possible part of a file is inspected. Check the coverage in each result.
- TXT
- Plain text and supported instruction signals.
- Markdown
- Text, comments and link destinations.
- HTML
- Parsed text, comments, selected attributes, metadata and supported inline hiding.
- DOCX
- Text, supported formatting and secondary document parts such as notes or comments.
- Extractable page text, annotations, form fields and metadata; attachments are inventoried.
- JSON
- Structured values and bounded adjacent-field context.
- CSV
- Textual cells and supported local context.
- TSV
- Textual cells and supported local context.
- XML
- Text, attributes and supported CDATA content.
- EML
- Message text, headers and supported HTML/preheader content.
- PPTX
- Extractable slide text and supported presentation channels.
- XLSX
- Cell text and supported workbook structure, including hidden sheets.
HIGH, INFO and NONE
HIGH points to a strong signal for review. INFO marks an unusual or contextual signal. NONE means no known pattern was found by the methods that ran. Counts are not a numerical risk score.
Why harmless content may be flagged
A quoted attack example, a security-training document or legitimately hidden editorial or accessibility text can resemble an attack. Context handling reduces some unnecessary alerts, but it cannot settle every case. Compare the excerpt with the original purpose.
When no finding appears
No known patterns found with the available methods does not prove the file safe. Review coverage for unavailable layers, especially image-only pages and unsupported embedded content.
Current limits
There is no OCR or image analysis. Some visual hiding, external or computed CSS, attachment contents and embedded objects are not interpreted. Language phrase coverage is incomplete, and the scanner does not test a downstream AI’s behavior.
Language coverage
All available language rule sets run on every scan independently of the interface language. Current phrase coverage includes English, German, Romanian, Mandarin Chinese, Hindi, Spanish, French, Arabic, Bengali, Russian, Portuguese, Urdu, Indonesian, Japanese, Nigerian Pidgin, Marathi, Telugu, Turkish, Tamil, Yue Chinese and Vietnamese. Structural signals can be language independent; phrase coverage is not exhaustive.
What to do with a finding
Inspect its location, compare the excerpt with the original context, and decide whether it is an instruction, an example or legitimate hidden content. Remove or isolate unexpected material and scan a revised copy when useful. Keep untrusted content separate from an AI agent’s governing instructions.