ChatGPT for Financial Services Moves the AI Fight Into the Data Contract

ChatGPT for Financial Services Moves the AI Fight Into the Data Contract

OpenAI bundles financial data into ChatGPT. The consequential shift is who controls evidence, entitlements and the path from a filing to a pitchbook.


A banker can lose an afternoon to a number that appears in three places and means something different in each. An earnings release reports an adjusted profit measure, a data terminal standardizes it, and an internal spreadsheet carries forward an older adjustment. The difficult work is not finding a number. It is deciding which number belongs in the valuation and preserving enough evidence for somebody else to challenge it.

That is the commercially important promise behind ChatGPT for Financial Services, which OpenAI announced on September 10, 2026. The product combines GPT-6 Astra with financial datasets hosted and indexed by OpenAI, source-level citations and tools for producing models, research notes and presentation materials. Morgan Stanley and Evercore helped shape it as design partners. This article, published September 12, examines that announcement rather than implying the product launched today. OpenAI's launch announcement is the primary source for the product claims.

The shift is more specific than a smarter chatbot arriving on Wall Street. OpenAI is moving into the layer between licensed information and finished financial work. That layer determines which sources analysts see, what evidence follows a figure into a spreadsheet, and which permissions survive when an answer becomes a client document. Better language-model reasoning helps. Control of that entire chain may matter more.

OpenAI is selling fewer handoffs, not just better answers

The launch describes three different routes to financial information. Some datasets are included and hosted by OpenAI. Some existing subscriptions are intended to become accessible through shared sign-in and entitlement integrations. Other information remains available through connectors. Those are distinct commercial and technical arrangements, not interchangeable versions of a search box.

OpenAI names Daloopa, PitchBook, LSEG News and Crunchbase among the built-in providers. It says customers can begin using the included datasets without negotiating separate contracts or configuring connectors. It separately says it is working with S&P Capital IQ, LSEG, MSCI, Dow Jones Factiva and Moody's on entitlement integrations. The latter wording describes ongoing work; it does not establish that every integration is already generally available. The launch page makes that distinction important.

For a procurement team, bundling could reduce the friction of an initial rollout. For an established bank, however, it creates a reconciliation exercise. Which datasets overlap with existing licenses? Which historical periods, fields and usage rights are included? Can an output be distributed outside the firm? A product announcement saying that data is included does not answer every redistribution, archive or audit question in an institution's contracts.

This is where the competitive contest becomes interesting. An AI interface that supplies both information and the document built from it can become the place where analysts spend their working day. The incumbent terminal does not have to disappear for its influence to weaken. It only has to become a background entitlement that a different interface invokes.

Four data brands do not make one homogeneous dataset

The providers serve different purposes. Daloopa emphasizes source-linked fundamentals and financial-model workflows. LSEG's financial news service combines news coverage with classification, tagging and delivery infrastructure. Crunchbase positions its product around private-company information and predictive signals. Those differences are visible in the providers' own descriptions, and they should remain visible inside any AI-generated analysis. Daloopa, LSEG and Crunchbase describe their products directly.

A reported financial result, a news report and a prediction about a private company's future financing are different kinds of evidence. Putting them into the same retrieval system must not flatten those distinctions. A model may use all three while preparing a company overview, but the reader needs to know which statement is historical, which is attributed reporting and which is an inference.

Consider a hypothetical acquisition screen. Public-company statements might support a margin calculation, private-market data might identify comparable transactions, and news coverage might reveal an announced restructuring. None automatically tells the analyst whether a restructuring charge is genuinely nonrecurring. That decision requires an explicit rationale. The assistant's useful role is to expose the supporting material and conflicting interpretations, not make the adjustment disappear behind a polished chart.

Data coverage also changes the meaning of absence. If a private company has no revenue figure in a licensed dataset, the assistant should not convert that gap into an estimate without saying so. An empty field is evidence about the dataset, not necessarily about the business. Buyers should deliberately include missing-data cases in their acceptance tests.

The adjusted EBITDA example is the right place to be skeptical

OpenAI uses profit-and-loss normalization as an example: a banker should be able to inspect the reconciliation and notes behind adjusted EBITDA, see excluded costs and decide how to use the measure in a valuation. That is a more meaningful demonstration target than asking a model for a generic company summary. It places the system in the messy space between reported facts and analytical judgment. OpenAI describes the workflow here.

A credible implementation needs to retain the original label, reporting period, units, currency and source location. It should distinguish management's adjustment from the analyst's adjustment. If an expense appears repeatedly, an analyst may decline to treat it as exceptional even when management does. The system should make that disagreement explicit rather than silently replacing one definition with another.

Financial information also has a preexisting machine-readable foundation. The SEC describes structured data as standardized pieces accessible to people and computers, while XBRL International explains how digital tags carry business meaning with reported facts. Those standards do not eliminate analytical judgment, but they give buyers a reason to ask whether an assistant preserves source semantics rather than treating every filing as undifferentiated text. The SEC's structured-data overview and XBRL International's explanation provide that context.

The same discipline applies when comparing companies. One issuer's adjusted earnings may exclude costs another issuer retains. A tidy peer table can therefore be less comparable than it looks. An assistant that rapidly fills the table but hides definitional differences accelerates an error. One that flags those differences before the analyst distributes the work actually improves the workflow.

For a pilot, the bank should give the assistant a small set of companies with known accounting complications, then compare its evidence trail against an analyst-reviewed reference. The evaluation should examine whether each adjustment can be reconstructed, not merely whether the final multiple is close. Two wrong adjustments can accidentally produce a plausible total. A final-answer score would miss that failure.

A citation is a starting point for an audit

OpenAI says the product can highlight specific tables and passages supporting figures and claims. That is valuable because the cost of checking an answer often determines whether checking happens. A link to an entire filing is much less useful than a pointer to the relevant table and footnote. But even granular citations do not establish that the reasoning between source and conclusion is correct.

The required evidence changes as work moves through the process. A research note may need a quotation and date. A valuation model needs assumptions, formulas and unit conversions. A pitchbook needs to show that its exported chart still reflects the approved workbook. Source retrieval, calculation and document generation each need their own checks. OpenAI's product description spans all three, which makes the boundaries especially consequential.

A useful internal rule would require every material figure to carry either a direct source or a clearly labeled derivation. That rule should survive export to Excel, Word and PowerPoint. Otherwise the evidence may be excellent inside ChatGPT and disappear in the file that a managing director actually sends. The launch describes firm templates and artifact generation, but it does not demonstrate every institution's downstream evidence-preservation requirements.

The following diagram describes an approval design a financial institution could adopt. It is an editorial recommendation, not a claim about OpenAI's internal implementation.

flowchart LR
    A[Licensed financial sources] --> B[Entitlement and period checks]
    B --> C[Source-linked working analysis]
    C --> D[Analyst adjustment review]
    D --> E[Approved workbook]
    E --> F[Firm template export]
    F --> G[Distribution and archive check]

Templates can standardize presentation and spread mistakes

Administrators can publish Excel, Word and PowerPoint templates through a dedicated admin page, according to OpenAI. That moves the product beyond answering questions toward producing recognizable firm deliverables. Templates can encode house style, standard sections and expected spreadsheet layouts. They can also make an unfinished analysis look deceptively finished. The announcement describes the template controls.

A familiar cover page carries institutional authority. Reviewers who would question a rough chatbot answer may skim a deck that resembles an approved pitchbook. The quality gate therefore needs to operate before branding is applied, or at least remain visible afterward. Draft status, missing evidence and unreviewed assumptions should not disappear because the assistant successfully follows the style guide.

A practical test is to provide a deliberately incomplete source package. Does the system leave an explicit gap, ask for the missing information, or fabricate a plausible-looking section to satisfy the template? That test is specific to financial artifact generation: the pressure to fill every cell and complete every slide can conflict directly with the duty to acknowledge uncertainty.

Template maintenance is another overlooked responsibility. If the firm updates a disclosure or changes a valuation convention, administrators need to know which generated artifacts used the old version. A product that centralizes templates creates an opportunity for governance, but the institution still needs version ownership, approval procedures and a way to trace exports back to the instructions in force when they were created.

Information barriers cannot be delegated to conversational politeness

OpenAI says the service builds on enterprise SAML single sign-on, SCIM provisioning and role-based access controls. It also describes role-level control over skills and apps, supported read and write actions, and multiple workspaces for information barriers. Those are relevant building blocks for financial institutions, where material nonpublic information and client confidentiality make broad access particularly dangerous. The security section is the vendor's stated position.

The institution still has to define the barrier. An analyst with access to a public research workspace should not receive a synthesized answer influenced by restricted deal materials. That remains true even if the output contains no direct quotation. The relevant question is whether retrieval and tool execution enforce the user's authorization before information reaches the model, not whether the model has been instructed to avoid mentioning secrets.

For a concrete pilot, create separate test workspaces representing public research and a restricted transaction team. Seed the restricted workspace with unmistakable synthetic information. Then test search, summaries, generated files and shared conversations from the public side. No actual client material is needed to expose an authorization failure. The test should include revoked access and copied artifacts, not just a first login.

OpenAI's broader enterprise privacy commitments say business data is not used for training by default and describe ownership and retention controls. Those commitments matter, but training policy is not the whole confidentiality problem. Access, logging, retention, deletion, export and third-party connections are separate questions. A reassuring answer to one should not be used as an answer to all.

Financial supervision follows the workflow into ChatGPT

FINRA's Regulatory Notice 24-09, published June 27, 2024, reminded member firms that existing rules and securities laws continue to apply when they use generative AI. The notice did not create a new exemption or a new set of requirements simply because the tool was an LLM. That historical guidance is relevant to the new product, not a September 2026 regulatory announcement. FINRA's notice states the boundary directly.

For a broker-dealer, supervision cannot stop at approving the software vendor. The firm must decide which tasks the system can perform, which outputs require review and how evidence is retained. A research draft, a client communication and an internal administrative summary have different consequences. Treating them as one undifferentiated category called AI use makes supervision less precise.

Marketing claims deserve similar care. The SEC's 2024 actions against Delphia and Global Predictions concerned false or misleading representations about their use of AI. They are not findings about OpenAI's new product. They illustrate why a financial institution should avoid promising clients that AI makes its analysis inherently more accurate unless it can substantiate that claim. The SEC's announcement explains those cases.

The bank's commercial language should match its tested workflow. Saying an assistant helps locate evidence is different from saying it independently validates an investment thesis. Saying it drafts a model is different from saying the model has been approved. The difference may feel pedantic during a product demonstration. It becomes essential when an investor or auditor asks what the institution actually represented.

Morgan Stanley's involvement needs a narrow reading

Morgan Stanley and Evercore are named as design partners for the new offering. That establishes their role in shaping it; it does not establish a firm-wide deployment, a quantified productivity result or an endorsement of every future feature. OpenAI says the early work focused on investment banking and equity research and that partner work will inform post-training and expansion. Those future-facing statements should stay future-facing.

Morgan Stanley has separately described an OpenAI-powered wealth-management tool, AI @ Morgan Stanley Debrief, that creates meeting notes with client consent and drafts follow-up material for an adviser to review. That earlier tool is useful context because it shows an institution defining a bounded workflow around consent and human review. It is not evidence that the new investment-banking product has already achieved the same operational maturity. Morgan Stanley's own announcement describes Debrief.

The distinction also prevents a misleading market narrative. A bank can use AI successfully in meeting administration while still facing unresolved problems in financial modeling. The source material, evaluation criteria and consequences differ. Product adoption should be tracked at the workflow level rather than inferred from a prestigious logo on a partner page.

For smaller firms, the partner list may still matter commercially. It signals which users influenced the initial design and which artifacts the product prioritizes. A private-credit team, an insurance analyst and an investment banker should not assume their needs are identical. The right buying question is which parts of the institution's actual work were represented in the design process.

The acceptance test should end outside the chat window

A serious pilot should follow one analysis from ingestion to distribution. Start with a known source package, require the assistant to identify the reporting periods and definitions, let it build a working model, and then export the result through the firm's template. Ask a reviewer who did not participate in the conversation to reconstruct the reasoning from the exported files alone.

That reviewer should be able to distinguish sourced values from assumptions, identify unresolved conflicts and locate the supporting material. If the reviewer must reopen a long chat and reverse-engineer the assistant's intent, the system has moved work rather than eliminated it. This is the operational meaning of OpenAI's promise to connect data, reasoning and artifact generation.

A second pilot should start from a changed fact: an amended filing, a corrected dataset or an updated assumption. The goal is to see whether the assistant updates every dependent output and clearly marks what changed. Financial work rarely ends after the first plausible answer. Revisions are where inconsistent periods, stale charts and broken provenance become expensive.

Risk-management frameworks can help organize such evaluations, but they do not replace the domain-specific tests. NIST's AI Risk Management Framework provides a broader governance reference. In this case, the meaningful measures are concrete: evidence coverage, unauthorized retrieval, formula integrity, correction propagation and the review time needed to approve an actual deliverable.

Restatements reveal whether the data layer has a memory of its own

There is another distinction a financial assistant must preserve: the difference between what is known now and what was available when a decision was made. A restated filing can improve the current dataset while making a historical investment memo impossible to reproduce if the original values are silently overwritten. The issue becomes more important when a hosted retrieval system supplies both the evidence and the finished analysis.

OpenAI's announcement emphasizes tracing figures across periods and interpreting annotations. A buyer should extend that promise into a point-in-time test. Give the system an original filing and a later correction, then ask two different questions: what is the latest reported result, and what information supported the earlier recommendation? Those questions should not receive the same undifferentiated answer. The product's stated research capabilities make this a relevant acceptance case.

The distinction is particularly sharp in event-driven research. An analyst studying the market's response to an earnings release needs the information that existed at the event, not a cleaned-up dataset assembled afterward. An assistant that retrieves the latest value without explaining the revision can introduce hindsight into an analysis that is supposed to describe a historical decision.

The bank should ask whether source versions, retrieval dates and correction notices remain accessible in the generated work product. If the answer depends on a separate archive, that archive needs to be part of the workflow rather than an emergency reconstruction exercise. The launch does not establish a complete point-in-time archive contract, so institutions should treat this as a question to verify rather than an assumed feature.

The Excel file is a calculation system, not a picture of one

A second practical distinction concerns the spreadsheet itself. OpenAI says the product can generate financial models and use firm-approved Excel templates. Buyers should determine whether a workbook contains inspectable formulas, explicit assumptions and clear dependencies, or merely values that happen to look like a model. Both can render attractively; only one may support the institution's intended revision process.

For example, a hypothetical acquisition model might include a revenue assumption, a margin assumption and a debt schedule. A reviewer should change one assumption and check whether the dependent outputs update coherently. The assistant should not need to regenerate the entire workbook just to preserve ordinary spreadsheet behavior. If some calculations are performed outside Excel, the handoff should make that clear and preserve enough information to reproduce them.

The review should also test sign conventions and units. Debt repayment, cash balances and expense adjustments can look plausible even when a sign has been reversed. A workbook can mix thousands and millions without producing an obvious visual warning. These are not exotic model failures; they are familiar spreadsheet risks that automated artifact generation can reproduce at greater speed.

OpenAI's template feature creates a place to enforce conventions, but it cannot decide the firm's analytical policy by itself. Finance teams should own the accepted formulas and disclosure language, while technology teams verify that generation preserves them. A template approved for one transaction type should not become an implicit license to improvise every other model in the same visual style.

That division of responsibility also makes vendor comparison more useful. Instead of asking competing assistants to produce an impressive deck, give each the same deliberately awkward source package and the same workbook requirements. Compare the resulting files after an analyst changes an assumption and another reviewer checks the evidence. The strongest product will remain understandable after the demonstration ends and the real work of revising the analysis begins.

The durable advantage is a defensible work product

OpenAI says ChatGPT for Financial Services is available to eligible financial institutions through its account and sales channels. The announcement does not provide a universal public price or prove that every data entitlement, geographic requirement and retention policy is supported for every buyer. Those remain procurement questions, not details an analyst should guess from a launch demonstration.

The opportunity is substantial because the product targets a real coordination cost. Analysts already move among data vendors, filings, spreadsheets, document editors and internal approval systems. Bringing those steps closer together could reduce repetitive work and make verification cheaper. But bundling also increases dependence on the intermediary that decides how sources are retrieved and how evidence is presented.

The winning implementation will not be the one that produces the prettiest pitchbook from the shortest prompt. It will be the one whose output survives an adversarial colleague asking where a number came from, why an adjustment was made, who was allowed to see the inputs and whether the exported file still says what the analyst approved. That is the test that turns an AI demonstration into financial infrastructure.

Subscribe to our newsletter

Get the latest posts delivered right to your inbox.

Subscribe on LinkedIn