Artificial intelligence
Data management
 •  
October 7, 2026

Building an AI control plane: Seven questions your AI architecture has to answer

Manvir Sandhu
By
Manvir Sandhu
Founder & Chief Innovation Officer

Part 2 of 2. The questions that show whether your AI architecture is ready to scale.

Treasury's Financial Services AI Risk Management Framework runs to 230 control objectives. It's a solid foundation, and so is the NIST AI Risk Management Framework it builds on. Both give your team the right terms and goals. Neither gives you a quick way to check whether the AI setup you already run actually works.

In Part 1, we covered why governance is the real constraint on AI at scale, and the three breakthroughs that put governed agentic AI within reach. These seven questions are the practical test. They'll show you where your AI architecture is solid and where the gaps are, so you know exactly where to start.

The seven questions at a glance

  1. Where is AI used, and by whom?
  2. Do we know every agent that exists, including the ones nobody registered?
  3. Who is asking, and on whose behalf?
  4. Is this action allowed?
  5. What did it cost?
  6. What may this agent see?
  7. What happened, and can we prove it?

1. Where is AI used, and by whom?

Start with the footprint. Which channels and applications does AI touch, and who meets it, whether that's employees, members, customers, or other systems?

Most institutions can answer this for deployments they sponsored. Fewer can answer it for AI that arrived inside a platform they already owned, like an assistant in the productivity suite or an agent feature in the CRM. The footprint is almost always wider than the project list.

Map it by channel and by audience. Flag anything that touches credit decisions, account servicing, or customer communications first, because examiners and fair lending reviews start there. Expect a single decision to cross several platforms. In Part 1, we followed one card dispute through the CRM, the data platform, the model provider, and the integration layer, and none of them saw the whole decision.

2. Do we know every agent that exists, including the ones nobody registered?

This takes two mechanisms. Inside out, your teams catalogue every agent as they create it, with its name, platform, owner, purpose, data access, risk tier, and go-live date. Outside in, you catch AI that never registers at the network edge and either block it or bring it under policy.

Most institutions attempt the first step and skip the second. That leaves the inventory only as complete as the honesty of the people filling it in.

The tooling for outside in now exists. MuleSoft's agent scanners connect to platforms like Amazon Bedrock, Google Vertex AI, and Agentforce, then add the agents they find to a central registry automatically. Databricks now registers agents and models in Unity Catalog alongside the data they use.

3. Who is asking, and on whose behalf?

Agents need identities the way employees do. Every request carries two of them, the agent's and the identity of the person or process it's acting for. Capture only the agent, and you can't say who authorized the action. Capture only the user, and you can't say which agent acted. Both identities, plus the delegation between them, make an action attributable.

Standards bodies are working on this now. NIST's National Cybersecurity Center of Excellence published a concept paper on AI agent identity and authorization in February 2026, and identity and privilege abuse is one of the ten risks in the OWASP Top 10 for Agentic Applications. Your identity team probably already manages service accounts. Agents need more than that, because a service account can't tell you whose request it's carrying.

The headless Salesforce pattern from Part 1 shows what good looks like. When an advisor asks Claude about a household, Salesforce checks the advisor's own profile, permission sets, and sharing rules. If the advisor isn't entitled to the data, Claude gets a denial with a reason.

4. Is this action allowed?

Policy has to be enforced at the moment of request. Does a person or a system verify, in real time, that this agent may take this action against this data for this person, and refuse when the answer is no?

Written policies and approval memos can't do that by themselves. Neither can guardrails that only filter prompts and responses. Enforcement needs a gateway that every agent call passes through.

Those gateways are now available on the platforms most institutions run. Databricks' Unity AI Gateway applies access policies and guardrails to model and agent calls at runtime. MuleSoft's agent kill switch revokes a misbehaving agent's tokens and credentials in one action, so the block holds at the identity layer as well as the network.

5. What did it cost?

AI is priced by consumption instead of by seat. That's unfamiliar territory for institutions whose entire buying history is per-user licensing.

An agent's cost depends on how often it runs, how much data it reasons over, and how expensive the underlying model is. All three are governance decisions, so they belong with your governance team as well as procurement. In Flexera's 2026 State of ITAM report, 59% of organizations said wasted AI spend increased year over year, and only 31% had accurate visibility into AI software. Consumption came up as a governance question in nearly every conversation we had at Dreamforce this year.

Treat AI cost as a renewal negotiation, and you'll find the real number in production, after your operations already depend on the agents. Set cost ceilings at intake, right next to the risk tier, and track spend by agent and by owner from day one. Gateways now enforce token budgets per agent, so a ceiling can hold in real time.

6. What may this agent see?

An agent is only as governed as the data underneath it. If sources are ungoverned, lineage is unclear, or quality is inconsistent, your request-layer policy ends up enforcing rules against data nobody can vouch for.

That's why data readiness keeps turning out to be the real constraint on AI speed, more than the model or the platform. We've written about how data maturity decides which AI projects survive. For agents, the question gets sharper. Can you say which data classes each agent can reach, whether any of it is nonpublic personal information, and whether it's accurate enough to act on? If not, governed, AI-ready data comes before the next agent.

The payoff is measurable. A $30 billion regional bank we partner with built use-case-scoped data products with a semantic layer, governed through Unity Catalog. Its agents now reach confident answers with fewer model calls and lower error rates than the pilot that ran without that layer. Part 1 covers how they did it.

The stakes are highest in lending. The EU AI Act classifies credit scoring as high-risk, though it doesn't bind US institutions. In the US, Regulation B still requires a specific reason when credit is denied. An agent that touches either needs data you can trace.

7. What happened, and can we prove it?

Every layer reports here. Observability is the real-time view of what agents are doing. Audit depth is whether you can reconstruct a single decision months later, including which agent acted, on whose behalf, against what data, under which policy version, and what a person did about it.

Many institutions have dashboards and mistake them for audit trails. A dashboard tells you last week's error rate, while an examiner wants to know about one loan application from March. GAO's 2025 review of AI in financial services points to data quality, biased lending decisions, and new cybersecurity threats as core risks, and each of those comes down to whether you can reconstruct what happened.

The regulatory picture raises the stakes. The April 2026 interagency model risk guidance, SR 26-2, places generative and agentic AI outside its scope while existing risk expectations still apply, so no checklist will tell you what's enough. Try the reconstruction test instead. Take one agent-assisted decision from six months ago and rebuild it end to end. Whatever you can't rebuild is the gap.

Answer the questions once, then reuse the answers

The institutions moving fastest answer these seven questions on their first deployment and reuse the answers on every one after. A $10 billion regional bank we work with built its agent registry, identity layer, cost controls, and audit trail alongside its first use case. Its governance board approved that pattern once, and every agent built to it has inherited the approval since.

Four tests for any AI control plane

Vendors now sell control planes, including Salesforce's new AI Control Plane. Once you know the questions, you can test whether any proposed control plane answers them across your real estate, beyond the one platform it ships with. Apply four tests:

  • Coverage. What does it actually see and control across every platform where your agents run?
  • Effectiveness. How well does it govern identity, policy, guardrails, and audit depth, beyond listing them as features?
  • Cost. What's the total cost at scale, including what it would cost to change course later?
  • Scalability. Does coverage hold as your agent fleet grows and your data foundation matures?

These tests keep the evaluation grounded. They also give your CISO, chief risk officer, and board risk committee a shared standard for judging readiness, one built on evidence rather than enthusiasm. That's usually what turns a governance conversation from a reason to say no into a path to yes.

Your environment won't look like the vendor demo

Vendor guides describe what their tools can govern under perfect conditions. No institution runs under perfect conditions. Most are working with:

  • A core system that's been around for decades
  • Integrations built by different teams over many years
  • Data that was never designed for AI agents to use

The only way to know what works is to test it in your own environment, on a real workflow, before you roll it out widely. Here's a simple way to start:

  1. Pick one AI agent that already works across more than one platform.
  2. Run it through all seven questions.
  3. Note where the answers are strong and where they're weak.

The strong answers give you a model to reuse. The weak ones show you what to fix next, and you'll find them before an examiner does. If you're earlier in the process, our Digital Maturity Assessment shows where your data and AI foundations stand today.

Get answers to all seven questions

Before you choose one platform to manage all your AI, make sure you know what's actually running across your institution and what each system can and can't control. In four to six weeks, Zennify's AI Control engagement gives you a clear picture of how your AI systems should fit together, a recommendation for where AI oversight should live, and a step-by-step roadmap. We build it around an AI use case already running at your institution, and it works with the oversight you already have, like model risk, vendor risk, and change management. 

Let's start the conversation.

$text$
$name$

$role$

Share this post
Facebook
LinkedIn