# Retrieval and ingestion ## Queries - Find every search or vector query whose results reach the context (for example `index.query`, `similarity_search`, `as_retriever`, `collection.query`). - Is there a filter based on the permissions of the people who will see the output? A query on a shared index with no permissions filter is a finding (LLM09). - Where does the identity used by the filter come from? It must come from the request, in code, not from a tool argument. - If the output goes to a group (a channel, a shared document), the filter must reflect what the whole group may see, not only the requester. - With several tenants, check that they are isolated: separate indexes or namespaces, or a tenant filter that the code always applies. ## Ingestion - Find the job that fills the index. Which sources does it read, and with which credentials? - Are permissions copied from the source into the metadata used by the filter? What does a document get when the source has no permissions set? It should be nobody, not everyone. Are permission changes and deletions in the source applied to the index? - Who can write to each source (LLM05)? Sources that many people can edit, such as wikis, shared drives and tickets, let any of them change what the model recommends. Report it for the owner to decide, even when it isn't a code issue. - Documents that contain attacker text by nature (incident reports, emails, support tickets, logs) are untrusted inputs: add them to the map. Confirm that the tool layer and the approval step don't depend on their content being safe. ## Results - Do results carry an identifier or a link to their source, so that the output can cite it (LLM07)? - Are the number and size of results limited before they enter the context (LLM06)?