Agencies sit in the middle of the AI accountability chain. Clients push obligations downward through master service agreements, questionnaires, and audit clauses, and the agency delivers work produced by employees, freelancers, and production partners across offices and time zones. The warranty in the contract is the agency's warranty regardless of whose account ran the generation, which means the agency needs the records regardless of who holds the login.
At agency scale, the failure mode is not the absence of a policy but the gap between policy and evidence. Most agencies now have an AI usage policy on paper. When a client audits, the question is not whether the policy exists but whether the stated process shows up in records: per-asset generation records, review sign-offs, and a match between what was declared and what was delivered. A policy document proves intent; only records prove practice.
New business raises the same question from the other side. AI sections have become standard in RFPs and procurement questionnaires, and an agency that can attach a redacted record from a real project answers them differently from an agency that attaches its policy PDF. Demonstrated practice is becoming a selection criterion, and it is one of the few compliance costs that doubles as a pitch asset.
One more thing has to be standardized alongside the records: what the agency claims about them. A record can show what existed, that it has not changed, and when it was recorded. It cannot show that a reference caused an output, and no agency-side record can show what a vendor trained a model on. Account teams that overpromise in a questionnaire create warranties the records cannot back, so brief every client-facing team on the boundary, which the provenance guide below sets out.
Where documentation breaks down
- Every account team documents differently, so answering one client's audit becomes a manual collection exercise across half the agency.
- Freelancers and production partners generate outside agency systems, and their records leave when the engagement ends.
- Legal signs MSAs with AI clauses that delivery teams never read, so contractual warranty and daily practice drift apart unnoticed.
- Pitch assets move into production with no record of how they were made in the pitch rush, and the gap surfaces at the worst possible moment.
- When a problem appears in one market's adaptation of an asset, tracing which office changed what takes days without a shared record chain.
A documentation routine that holds up
- Put one record standard in the production handbook: which fields, captured when, stored where, identical across accounts and offices.
- Extend it contractually to freelancers and production partners, so a delivery is not complete until the record arrives with the files.
- Have account leads reconcile records against deliverables at each milestone, instead of discovering the gaps during a year-end audit.
- Keep client-facing summaries separate from the internal record, so answering one client's audit never means exposing another client's work.
- Brief legal and new business on what the records can and cannot support, so questionnaire answers and contract warranties stay inside the evidence.
Frequently asked questions
A client audit asks us to evidence our AI policy. What counts as evidence?
Records of practice: per-asset generation records, review sign-offs, and a reconciliation showing that what was declared matches what was delivered. The policy document itself only shows intent, and auditors know the difference.
Our subcontractor generated the asset. Whose record is it?
Contractually, the agency answers to the client, so the record has to reach the agency as part of the subcontractor's delivery. Make that a term of the engagement rather than a favor to ask later, because by the time you need it the engagement is usually over.
Do internal enterprise AI tools need records too?
Yes. The client's question is about the work, not about which category of tool produced it. Anything generative that shaped a deliverable belongs in the record, whether it was a public tool, an enterprise suite, or something built in-house.
How should we answer training-data questions in RFPs?
Honestly and narrowly. No agency-side record can show what a vendor's model was trained on, and an RFP answer that implies otherwise becomes a warranty you cannot back. What you can evidence is your own process and your tool choices, and a well-kept record makes that answer strong enough on its own.
Where does the EU AI Act fit into all of this?
Read the Article 50 guide for what the transparency obligations actually say. The practical point for agencies is that the record fields those obligations imply overlap almost entirely with what clients already ask for, so one record standard answers both instead of two systems answering one each.