Harvey’s GPT-6 Astra Workflow Shows Why Legal AI Must Carry Its Context

Harvey’s GPT-6 Astra Workflow Shows Why Legal AI Must Carry Its Context

Harvey says GPT-6 Astra improves legal drafting by preserving context, but dependable legal AI depends on provenance, review, and the limits of confidence.


A legal draft can be grammatically perfect and still be unusable if nobody can explain which clause, precedent, or client instruction shaped it. Harvey says its GPT-6 Astra workflow is designed to turn legal context into stronger drafts. The significant story is not that a model writes faster. It is whether the system can preserve the evidence trail that makes a lawyer willing to sign the work.

flowchart LR
R[Reported claim] --> E[Evidence boundary] --> D[Deployment decision] --> O[Observable outcome]

The draft is only the visible layer

Harvey’s announcement presents GPT-6 Astra as a way to move from raw legal context to more confident drafting. That wording is a product claim, not an independent finding, and the relevant question is which parts of the context the system can retrieve and show to its reviewer.

A matter team can trust this workflow only when every important sentence has an identifiable source, an applicable date, and a reviewer who can reject the inference.

The legal test is not verbal polish. It is whether a second lawyer can reproduce the path from client instruction to retrieved authority to approved language without relying on the model’s memory. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

Evidence that should travel with the claim

A useful audit record would show the matter boundary, the documents selected, the conflicts found, and the point at which professional judgment replaced generation. The direct announcement supplies the vendor’s stated capability and date; the institutional sources provide comparison points rather than independent validation.

Context is a legal control

Legal context is layered: the engagement letter, the client’s commercial goal, governing law, procedural posture, internal playbooks, and the latest correspondence can all constrain a sentence. A model that sees only the visible document may produce elegant language that violates an instruction stored elsewhere.

The legal test is not verbal polish. It is whether a second lawyer can reproduce the path from client instruction to retrieved authority to approved language without relying on the model’s memory.

A useful audit record would show the matter boundary, the documents selected, the conflicts found, and the point at which professional judgment replaced generation. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

Confidence needs a source trail

Confidence is useful only when it is attached to inspectable support. A reviewer should be able to distinguish a quoted authority from a model inference, a client fact from a placeholder, and a proposed argument from a conclusion already approved by counsel.

A useful audit record would show the matter boundary, the documents selected, the conflicts found, and the point at which professional judgment replaced generation.

That makes context a control surface: changing the retrieved material should change the explanation, while changing the typography should not change the underlying claim. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

Why retrieval changes the review job

Context retrieval changes review from proofreading into evidence inspection. The lawyer is no longer checking only whether a paragraph sounds right; the lawyer is checking whether the system selected the right materials and preserved the conditions under which those materials apply.

That makes context a control surface: changing the retrieved material should change the explanation, while changing the typography should not change the underlying claim.

A matter team can trust this workflow only when every important sentence has an identifiable source, an applicable date, and a reviewer who can reject the inference. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

Evidence that should travel with the claim

The legal test is not verbal polish. It is whether a second lawyer can reproduce the path from client instruction to retrieved authority to approved language without relying on the model’s memory. The direct announcement supplies the vendor’s stated capability and date; the institutional sources provide comparison points rather than independent validation.

The difference between assistance and advice

Legal assistance becomes advice when the product or user treats a generated position as a conclusion without professional judgment. A workflow can draft, compare, summarize, and expose missing evidence while leaving the legal determination with an accountable practitioner.

A matter team can trust this workflow only when every important sentence has an identifiable source, an applicable date, and a reviewer who can reject the inference.

The legal test is not verbal polish. It is whether a second lawyer can reproduce the path from client instruction to retrieved authority to approved language without relying on the model’s memory. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

Astra’s vendor claim needs a boundary

The announcement establishes Harvey’s positioning and the availability of the workflow. It does not establish that Astra is more accurate on every matter, that every jurisdiction is covered, or that a customer’s confidential material is handled in an identical way across products.

The legal test is not verbal polish. It is whether a second lawyer can reproduce the path from client instruction to retrieved authority to approved language without relying on the model’s memory.

A useful audit record would show the matter boundary, the documents selected, the conflicts found, and the point at which professional judgment replaced generation. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

Long documents create selective memory

Long matters create a selective-memory problem. A model may have access to thousands of pages while the answer is shaped by a small retrieved slice. The system must show enough surrounding context for a reviewer to notice when an exception or later amendment changes the meaning.

A useful audit record would show the matter boundary, the documents selected, the conflicts found, and the point at which professional judgment replaced generation.

That makes context a control surface: changing the retrieved material should change the explanation, while changing the typography should not change the underlying claim. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

Evidence that should travel with the claim

A matter team can trust this workflow only when every important sentence has an identifiable source, an applicable date, and a reviewer who can reject the inference. The direct announcement supplies the vendor’s stated capability and date; the institutional sources provide comparison points rather than independent validation.

The client file is not a prompt

A client file is not a neutral prompt. It contains privileged material, negotiation strategy, personal information, and documents written for different audiences. Treating all text as equally authoritative can make a stale email look like a current instruction.

That makes context a control surface: changing the retrieved material should change the explanation, while changing the typography should not change the underlying claim.

A matter team can trust this workflow only when every important sentence has an identifiable source, an applicable date, and a reviewer who can reject the inference. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

Citation quality is part of quality

Citation quality has two dimensions: whether a source is present and whether it actually supports the proposition. Legal teams should test both, because a link to the right case can still accompany a claim that the case does not make.

A matter team can trust this workflow only when every important sentence has an identifiable source, an applicable date, and a reviewer who can reject the inference.

The legal test is not verbal polish. It is whether a second lawyer can reproduce the path from client instruction to retrieved authority to approved language without relying on the model’s memory. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

The human reviewer needs the right interface

A reviewer interface should surface assumptions, retrieved passages, unresolved conflicts, and the requested action. Hiding those elements behind a polished final draft transfers the hardest work to the person with the least time to inspect it.

The legal test is not verbal polish. It is whether a second lawyer can reproduce the path from client instruction to retrieved authority to approved language without relying on the model’s memory.

A useful audit record would show the matter boundary, the documents selected, the conflicts found, and the point at which professional judgment replaced generation. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

Evidence that should travel with the claim

That makes context a control surface: changing the retrieved material should change the explanation, while changing the typography should not change the underlying claim. The direct announcement supplies the vendor’s stated capability and date; the institutional sources provide comparison points rather than independent validation.

Confidentiality follows the data path

Confidentiality follows the data path through ingestion, indexing, retrieval, model inference, logging, exports, and deletion. A contract that promises secure inference does not answer every question about derivative indexes or human access to debugging records.

A useful audit record would show the matter boundary, the documents selected, the conflicts found, and the point at which professional judgment replaced generation.

That makes context a control surface: changing the retrieved material should change the explanation, while changing the typography should not change the underlying claim. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

Automation must stop at authority

Authority must be explicit. Drafting a motion, sending it to a client, filing it, and changing a matter record are different actions with different approval requirements. A useful agent keeps those boundaries visible instead of treating them as one continuous task.

That makes context a control surface: changing the retrieved material should change the explanation, while changing the typography should not change the underlying claim.

A matter team can trust this workflow only when every important sentence has an identifiable source, an applicable date, and a reviewer who can reject the inference. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

What a defensible pilot measures

A defensible pilot should measure source selection, unsupported assertions, correction time, missed instructions, and reviewer agreement. Counting documents drafted or hours saved alone can reward a system that produces more work for the final reviewer.

A matter team can trust this workflow only when every important sentence has an identifiable source, an applicable date, and a reviewer who can reject the inference.

The legal test is not verbal polish. It is whether a second lawyer can reproduce the path from client instruction to retrieved authority to approved language without relying on the model’s memory. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

Evidence that should travel with the claim

A useful audit record would show the matter boundary, the documents selected, the conflicts found, and the point at which professional judgment replaced generation. The direct announcement supplies the vendor’s stated capability and date; the institutional sources provide comparison points rather than independent validation.

The economics are more than tokens

The economics include review, integration, retention, access controls, and the cost of a wrong legal position. A faster first draft can lose its value if each paragraph requires a second lawyer to reconstruct where it came from.

The legal test is not verbal polish. It is whether a second lawyer can reproduce the path from client instruction to retrieved authority to approved language without relying on the model’s memory.

A useful audit record would show the matter boundary, the documents selected, the conflicts found, and the point at which professional judgment replaced generation. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

The portability question

Legal teams should avoid tying their matter data to one model’s hidden conventions. Portable source identifiers, structured matter metadata, and evaluation sets make it possible to compare providers without starting the governance process again.

A useful audit record would show the matter boundary, the documents selected, the conflicts found, and the point at which professional judgment replaced generation.

That makes context a control surface: changing the retrieved material should change the explanation, while changing the typography should not change the underlying claim. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

What Harvey still has to prove

Harvey’s workflow points toward a useful direction: context should make a draft more accountable, not merely more fluent. The remaining proof is operational—whether firms can reproduce the reasoning path, correct it quickly, and show a client why the final work deserves trust.

That makes context a control surface: changing the retrieved material should change the explanation, while changing the typography should not change the underlying claim.

A matter team can trust this workflow only when every important sentence has an identifiable source, an applicable date, and a reviewer who can reject the inference. The claim should therefore be paired with an observable review event, a documented exception path, and a clear owner for correction.

Evidence that should travel with the claim

The legal test is not verbal polish. It is whether a second lawyer can reproduce the path from client instruction to retrieved authority to approved language without relying on the model’s memory. The direct announcement supplies the vendor’s stated capability and date; the institutional sources provide comparison points rather than independent validation.

The file should explain itself

A legal AI system earns its place when the matter file remains intelligible to someone who did not watch the original interaction. That means preserving the source passages, the document dates, the instruction hierarchy, the reviewer’s edits, and the reason a disputed sentence survived. It also means distinguishing generated work from approved work in exports and client communications. A firm should be able to answer a simple question months later: why was this language used, and what evidence was available at the time?

The best pilot will therefore compare not only first-draft speed but review reconstruction. Give two teams the same matter, ask them to find unsupported claims, and measure how long it takes to locate the governing material. Test a later amendment, a contradictory client email, and a document with a misleading heading. Those cases reveal whether the system understands legal context or merely retrieves text that resembles the request.

Harvey’s Astra story is consequential because it moves the conversation from generic writing assistance toward work product with professional consequences. That move should raise the evidence standard. A strong product can still be useful without claiming autonomy; it can make the lawyer faster while making the lawyer’s reasoning more visible. The boundary is not a weakness. It is what allows a firm to adopt the tool without pretending that fluency has replaced responsibility.

A second control is version discipline. The firm should retain the model identifier, retrieval configuration, prompt policy, and source snapshot used for a draft. Otherwise a later replay may produce different wording and make it impossible to tell whether the lawyer’s decision, the evidence, or the system changed. Version discipline protects the client and gives the vendor a fair way to investigate defects.

A matter team can also compare the system against ordinary search and a human research baseline. The question is not whether the model produces a longer draft, but whether it finds the controlling material with fewer missed exceptions and less duplicated effort. That comparison keeps the evaluation grounded in the work the firm actually wants to improve.

The baseline should include the time spent correcting citations and checking privilege boundaries. Those costs are part of the system’s actual performance, even when they do not appear in a product demonstration.

That baseline should be repeated across jurisdictions and matter types, because a result on one kind of contract does not establish performance on litigation, employment, or regulatory work.

A firm should also test whether the workflow changes the distribution of review work. If senior lawyers spend less time searching but more time correcting subtle omissions, the benefit may be real but different from the marketing claim. Matter economics should show that tradeoff rather than hide it inside a single productivity number.

Clients will increasingly ask how AI contributed to their work. A defensible answer is not that the model was accurate; it is that the firm used approved sources, retained a review record, limited the system’s authority, and corrected errors through a known process. That is a stronger promise because it can be inspected.

Sources and reporting trail

The article distinguishes announcement dates from independent verification. Direct primary and institutional sources reviewed for the factual claims and limits include:

Subscribe to our newsletter

Get the latest posts delivered right to your inbox.

Subscribe on LinkedIn