Your firm has made the big call. Loan files, SAR case files, and exam reports will stay on systems you control, and AI will run there too. What remains is practical: what has to be built, who will run it, and how long it takes before a loan officer can ask it a question.
Leadership will ask three questions along the way. Where does our data go? Who can see it? What do we tell the examiner? A plan that answers all three before the build starts tends to clear review faster.
Which data has to stay inside?
Start with the data, before anyone sizes a server. For each job you want AI to do, ask what it would cost if that data leaked tomorrow. The answer tells you which sources go into the private system first, and which can stay in the tools you already use.
| The data | What a leak could cost | Where it should run |
|---|---|---|
| Loan files, account records, and other customer data. SAR and AML case files. Exam reports. A fund’s research notes and trade plans. | Notices to customers, an exam finding, a breach of SAR confidentiality, or an MNPI problem. | On your own servers. |
| Procedures, board decks, vendor contracts, and draft policies. | A lost edge or a bad week, rarely a legal event. | Your call. |
| Marketing copy, published rate sheets, and public filings. | Little or nothing. | Any tool you already trust. |
The top row carries legal duties on top of business risk. A SAR, and any information that would reveal one, is confidential by federal rule.[1] Customer data falls under the safeguards rules your regulator enforces: the banking agencies’ security guidelines for banks, the NCUA’s guidelines for credit unions,[2] and Reg S-P for brokers, funds, and advisers.[3] Advisers also need written policies to keep material nonpublic information from being misused.[4]
The middle row takes judgment. Procedures and board decks show how your firm works, and a rival would value them. Keeping them with the top row gives you one system to govern.
Where can the servers run?
You have three places to put the servers and two ways to run them. The places are a server room you already own, rented space in a colocation data center, or your firm’s own account with a cloud provider. Your IT team can run any of them, or your MSP can.
- Your own server room. A fit if you already run core systems in-house and have the power and cooling. You buy the GPU servers, and the order is the long wait.
- Colocation. You own the servers and rent rack space, power, and cooling in another company’s data center. A fit if you have no server room to spare.
- Your own cloud account. GPU servers rented inside your firm’s account, under your own network rules. It starts fastest, since there is no hardware to ship. The cloud provider joins your vendor list if it isn’t on it already.
Placement also shapes the bill. Owned servers are a one-time purchase plus power, cooling, and support. A cloud account bills for the GPU servers you keep running. Either way the bill follows the servers, so staff can ask as many questions as the servers can handle.
What goes into the stack?
A working stack has five parts. The first is the GPU servers and the software that serves the model. The other four turn a model into something your firm can use and defend in an exam: access and audit, data connections, process design, and testing. Build those four in that order.
1. GPUs and model serving. The model runs on GPU servers. A serving engine such as vLLM loads an open-weight model, for example one from Meta’s Llama family, and answers requests. Size depends on how many people use it at once and what they ask. A short answer from a policy manual takes far less GPU time than a summary of a long loan file, or a nightly job that reads every new file. Pick the model by testing it on your own documents, and plan how you’ll swap in a better one later.
2. Access and audit. Staff sign in through your existing single sign-on, over SAML or OIDC, and each answer follows the permissions that person already has. A loan officer’s question should reach only the files that loan officer can open. Every action by a person or an agent goes into a log your compliance team can search when an exam request arrives. Settle early who counts as the user when an agent acts on its own, such as a job that runs overnight.
3. Data connections. Connectors pull from the systems that hold your files: SharePoint, file shares, the loan origination system, and the document system. Each source raises its own questions. How often does content sync? Which version of a policy is current? What happens to scanned PDFs and tables buried in attachments? Search has to return the passage behind each answer, so the answer can cite it. The steps that read, split, and index files have to run inside too. If one of them sends text to an outside service to build search data, your files have left, whatever the chat window shows.
4. Process design. Procedures tend to be spread across PDFs, intranet pages, and training decks. Loading them turns up conflicts, like two versions of the wire procedure with different callback rules, and someone has to settle those before the AI cites either one. Then decide which steps the AI drafts, where a person approves, and how exceptions get recorded.
5. Testing and upkeep. Build a test set from real questions and real files, including the hard ones: the loan exception, the policy that changed last quarter, the question your files can’t answer. Agree on what counts as a pass, how often someone reviews answers, and what triggers a rollback. Then plan upgrades, with a named person to test each new model, approve it, and undo it if needed.
Access and audit go first because data connections mean little until you know who may see what. Process design needs stable data, and testing needs a settled process.
Running inside your own walls lets much of this reuse what you have. Identity comes from your directory. Logs can feed the SIEM your security team already watches. Procedures can be indexed as they are, since they stay inside.
That is the list your team would build. Lumen, the private AI platform from Cognetryx, comes with those parts already built and working together, and it runs the AI model you choose. The Why Cognetryx page sets your build list next to what Lumen already includes.
How long does it take?
Lumen arrives with the five parts already built, so the date depends on your hardware and your firm’s decisions. A team that builds its own stack waits on the same hardware, plus the work of building the five parts.
These decisions move the date, roughly in the order they come up:
- Where it runs. Your own hardware waits for the order to ship. A cloud account can start on rented GPUs.
- The first sources. Start with one or two collections, such as the loan policy manual or the BSA procedures, and add the rest once those work.
- The access map. IT, compliance, and each business owner sign off on who can see which collections.
- Security review and vendor due diligence. Start both the week the hardware is ordered, so they finish during the wait.
- The test set and pass line. Write them before launch, so go-live day has a clear yes or no.
Settle these while the hardware is on order. Vague answers early turn into delays later, and the GPU wait is time you can use.
How does it change third-party risk?
It changes who touches your data, and every party that remains still counts. The 2023 interagency guidance, issued by the FDIC as FIL-29-2023, covers any business arrangement between a bank and another company, with or without a contract. It doesn’t mention AI. A software vendor, the MSP that runs your servers, and your cloud provider all fall under it.[5]
The guidance, from the Fed, FDIC, and OCC, takes each relationship through a life cycle: planning, due diligence, contract negotiation, ongoing monitoring, and termination. It also scales the work to the risk. A vendor whose software your team installs is a different risk from one that keeps your customer files on its own servers.
Running the model inside your walls cuts the number of outside parties that hold customer data. Rate each party that remains by what it can reach. An MSP with admin rights on the AI servers ranks high. The company that sold you the hardware ranks low. If the software vendor can log in for support, write down when, how, and who approves it.
On September 11, 2026, the Fed, FDIC, OCC, and NCUA proposed guidance to help banks and credit unions match third-party work to the risk of each relationship. Once final, it would replace the 2023 version.[6] For now, credit unions have the NCUA’s 2007 letter on third-party relationships, which covers planning, due diligence, and controls.[7]
Brokers, funds, and advisers answer to the SEC’s Reg S-P. Its 2024 amendments require firms to oversee service providers and to make sure each one reports a breach of a customer data system it runs within 72 hours.[3]
Who runs it after go-live?
Split the work three ways and name the people before launch. IT, or your MSP, runs the servers, patches, model updates, and access settings. Each business unit owns its files and approves what the AI can read. Compliance reviews the logs and keeps the record examiners will ask for.
Day to day, it looks more like running a database than running a software team. New policies go into the index. Updates get tested and installed. Odd answers get checked. Each job is small, and each needs a name next to it.
If your IT team is two people, your MSP can run the AI servers the way it runs the rest of your network. Add them to its contract and to the monitoring your vendor program already does. Cognetryx ships Lumen updates as signed packages your IT team installs on its own schedule, so each upgrade can go through your normal change control.
What should you ask a vendor?
Ask the same questions whether you’re reviewing a platform or your own team’s build plan. Each one maps to a part of the stack above. When an answer comes back vague, expect that part to show up later as a delay or an exam finding.
- Does anything leave? Prompts, answers, logs, usage stats, support data, and the file prep steps all count. “Only for monitoring” still means it leaves.
- Does it use your sign-in and your permissions? Have them show a user asking about a file that user can’t open.
- Can a reviewer check each answer? Click a citation and see whether it opens the right passage in the current version. The guide to checking AI answers covers the rest.
- What gets logged, and who can search it? Ask to see a real log entry for an agent run, and how long logs are kept.
- What proof shows it works on your files? Ask for a test on your own documents and terms, with the hard cases in it.
- How do updates arrive, and who can roll one back? Find out how fast a bad update can be undone.
- How do retention and deletion work? You should be able to keep a chat as a record, or delete it when your policy says so.
- Where does a person still have to approve? Write it down. For broker-dealers, FINRA’s Regulatory Notice 24-09 says its rules, supervision included, apply to generative AI just as they apply to any other tool.[8]
- What does it cost at real use? A meter that grows with every question, or a fixed line you can budget.
Put the same questions to us. In Lumen, answers cite their source, and a click opens the document with the passage highlighted. Answers follow each person’s existing permissions, with SSO over SAML or OIDC. The Compliance Portal logs chats, tool calls, agent runs, approvals, deletion requests, exports, sign-ins, and permission and config changes, and your compliance team can search it. The platform has one fixed price, quoted after a short scoping call.
See the stack you’d build, ready to deploy
Bring your build list and vendor questions to a demo, and check each one against Lumen.
See it in actionSources
- 31 CFR 1020.320(e), confidentiality of SARs filed by banks. A SAR, and any information that would reveal one, is confidential and may be disclosed only as the rule allows. law.cornell.edu
- Interagency Guidelines Establishing Information Security Standards, 12 CFR part 364, appendix B (the FDIC’s version), and NCUA, Guidelines for Safeguarding Member Information, 12 CFR part 748, appendix A, which sets standards under the Gramm-Leach-Bliley Act. Both include a section on overseeing service providers. 12 CFR 364; 12 CFR 748
- SEC, amendments to Regulation S-P, adopted May 15, 2024 (SEC press release 2024-58). Reg S-P covers broker-dealers, investment companies, registered investment advisers, and transfer agents. The amendments add an incident response program that includes oversight of service providers. The 72-hour notice from service providers is at 17 CFR 248.30(a)(5)(i)(B). sec.gov; 17 CFR 248.30
- Investment Advisers Act of 1940, section 204A (15 U.S.C. 80b-4a), which requires advisers to keep written policies to prevent misuse of material nonpublic information. law.cornell.edu
- Federal Reserve, FDIC, and OCC, Interagency Guidance on Third-Party Relationships: Risk Management, final June 6, 2023, issued by the FDIC as FIL-29-2023. It addresses “any business arrangement between a banking organization and another entity, by contract or otherwise,” sets out a five-stage life cycle, and ties the depth of oversight to the risk. The guidance does not mention AI. fdic.gov
- FDIC, FIL-58-2026, Proposed Interagency Third-Party Risk Management Guidance, September 11, 2026, issued with the Federal Reserve, the NCUA, and the OCC (OCC Bulletin 2026-46). Any finalized guidance would replace the 2023 guidance. fdic.gov; occ.gov
- NCUA, Letter to Credit Unions 07-CU-13, Evaluating Third Party Relationships, December 2007, with Supervisory Letter 07-01. ncua.gov
- FINRA, Regulatory Notice 24-09, on member obligations when using generative AI and large language models, June 27, 2024. finra.org
This article is informational and not legal or compliance advice. Confirm how any rule applies to your firm with your own counsel and compliance team.