Where would AI save you the most time?Free workbook: find the tasks worth automating.
Get the workbook

Home/Articles

AISecurity

AI Data Exposure Is a Security Problem Nobody's Talking About

By Al Bunch · · 5 min read

If you’re adding AI features to an existing application, there’s a security problem you’re probably not thinking about: what data are you sending to the LLM?

Most tutorials and demos take the same approach — dump your database schema into the prompt so the AI can generate SQL, answer questions, or build reports. It works. It’s also reckless.

The Problem

Here’s what a typical AI integration looks like when someone wants to let users query their data with natural language:

// The naive approach
$schema = $this->getFullDatabaseSchema();
$prompt = "Given this schema:\n{$schema}\n\nWrite SQL for: {$userQuestion}";
$response = $openai->chat($prompt);

That getFullDatabaseSchema() call just sent everything to the LLM. Every table. Every column. Including:

  • users.password — hashed, but still
  • users.api_token — now the AI knows your token column name and type
  • users.ssn — if you store it
  • payments.stripe_customer_id — internal billing identifiers
  • internal_flags.is_shadow_banned — operational fields users should never know exist

The AI doesn’t need any of this to answer “how many form submissions did we get last month?” But it can see all of it. And it will sometimes reference these fields in generated SQL if the question is vague enough.

Even if the AI never leaks this data directly to the user, you’re still sending it to a third-party API in every prompt. That’s a compliance issue for anyone handling healthcare data, financial data, or PII.

The Common Fix That Doesn’t Work

The first instinct is to hardcode a “safe” schema:

$safeSchema = "
    forms: id, name, created_at
    sessions: id, form_id, browser, city, country, submitted_at
";

This works until it doesn’t. You add a column to sessions, forget to update the hardcoded string, and now the AI is working with a stale schema. Or worse — a new developer adds a table to the safe schema without realizing it contains sensitive fields.

Hardcoded schemas drift. They always drift.

A Better Approach: Opt-In Exposure

When we built our natural language query system, we wanted the security decision to live where the data is defined — on the entity itself. Not in a config file, not in a service class, not in a hardcoded string.

We created a custom PHP attribute called #[AiSafe]:

#[Attribute(Attribute::TARGET_PROPERTY)]
class AiSafe
{
    public function __construct(
        public ?string $description = null
    ) {}
}

Then on the entity:

#[ORM\Entity]
class Session
{
    #[ORM\Id]
    #[ORM\GeneratedValue]
    #[AiSafe(description: 'Unique session identifier')]
    private int $id;

    #[AiSafe(description: 'Web browser used')]
    #[ORM\Column(length: 100)]
    private string $browser;

    #[AiSafe(description: 'City from geolocation')]
    #[ORM\Column(length: 100, nullable: true)]
    private ?string $city;

    // This field exists but the AI will never know about it
    #[ORM\Column(length: 45)]
    private string $ipAddress;

    // Neither will this one
    #[ORM\Column(type: 'json', nullable: true)]
    private ?array $rawHeaders;
}

The schema extractor reads Doctrine metadata and only includes fields marked with #[AiSafe]:

class SchemaExtractorService
{
    public function extractSafeSchema(): string
    {
        $schema = [];

        foreach ($this->entityManager->getMetadataFactory()->getAllMetadata() as $meta) {
            $reflection = $meta->getReflectionClass();
            $fields = [];

            foreach ($meta->getFieldNames() as $field) {
                $prop = $reflection->getProperty($field);
                $attrs = $prop->getAttributes(AiSafe::class);

                if (!empty($attrs)) {
                    $aiSafe = $attrs[0]->newInstance();
                    $type = $meta->getTypeOfField($field);
                    $desc = $aiSafe->description ? " -- {$aiSafe->description}" : '';
                    $fields[] = "  {$field} ({$type}){$desc}";
                }
            }

            if (!empty($fields)) {
                $table = $meta->getTableName();
                $schema[] = "{$table}:\n" . implode("\n", $fields);
            }
        }

        return implode("\n\n", $schema);
    }
}

What the AI sees:

sessions:
  id (integer) -- Unique session identifier
  browser (string) -- Web browser used
  city (string) -- City from geolocation
  country (string) -- Country from geolocation
  submitted_at (datetime) -- When the form was submitted

What the AI doesn’t see: ipAddress, rawHeaders, and anything else not explicitly marked. Those fields are invisible — not hidden, not filtered, not present at all.

Why This Works

It’s declarative. The security decision lives on the property itself, right next to the ORM mapping. A developer adding a new column has to actively choose to make it visible to the AI.

It fails safe. If you add a field and forget to add #[AiSafe], it’s excluded. The default is invisible. This is the opposite of most security models where you have to remember to restrict access.

It’s auditable. Want to know exactly what the AI can see across your entire application? grep -r "AiSafe" src/Entity/ — done. No chasing config files or service methods.

It doesn’t drift. The schema the AI sees is always in sync with your actual entities because it’s extracted from the same source of truth — Doctrine metadata. Add a column, mark it safe, it appears. Remove a column, it’s gone.

The description field is documentation. The optional description parameter on #[AiSafe] serves double duty — it gives the AI context about what the field means (improving query accuracy) and it documents the field’s purpose for other developers.

Defense in Depth

Controlling what the AI sees is one layer. If the generated SQL actually runs, add two more:

  • Run it as a read-only database user that can only SELECT from the tables and columns you’ve exposed. Even a malicious or confused query can’t change data.
  • Scope every query to the current customer. In a multi-tenant app, apply the tenant filter in code or with row-level security, never by trusting the AI to add a WHERE clause.

The attribute decides what the model knows about. The database role decides what a query can touch. You want both.

The Principle

Opt-in exposure, not opt-out hiding.

Every AI integration that touches your data should answer one question clearly: what exactly can the AI see? If you can’t answer that by looking at your entity definitions, you have a security gap.

The #[AiSafe] pattern isn’t complicated. It’s a simple PHP attribute, a schema extractor that reads it, and a principle that unmarked fields don’t exist as far as the AI is concerned. You could implement this in an afternoon.

The hard part isn’t the code. The hard part is remembering to think about it at all.


We built this pattern for a natural language query system: users type questions in plain English, the AI writes SQL, and results render as live charts. If you’re adding AI to your application and thinking about what it should and shouldn’t see, see our AI automation services or let’s talk.

Free download

The AI Automation Workbook

Find the tasks worth handing to AI, estimate the payback, and plan a pilot you can measure.

Get the free PDF

Ready to Get Your Systems Connected?

Tell us what tools you're using and what's not working.