Operating question
Useful AI often becomes quieter as it improves: it applies an expert's criteria, brings context into view without taking the floor, and preserves the right working material rather than simply consuming more tokens.
Human-Centered Architecture
3 Things AI: The “Help Without Taking Over” Edition
For
Leaders and workflow owners
You will leave with
3 operating decisions
Reading mode
5 min · 3 verified sources
Reading guide3 decisions · 4 sections+
Decision points
- 01Write down an expert's review questions before asking AI to apply them.
- 02Test small, source-linked meeting prompts before adding live listening.
- 03Measure the context an agent receives and uses, not only its token limit.
Three fresh signals show how to reuse expert judgment, support a meeting without interrupting it, and measure what an AI agent actually remembers.
The best AI improvement this week may be a quieter one: give people the right help without making the tool the centre of the work.
That means capturing judgment before automating it, protecting attention while a conversation is happening, and checking what an agent actually receives rather than trusting a large context number.
1. Turn one expert's review into a useful first pass
A proposal reaches the same senior advisor every Friday. She checks whether the claim is supported, the next step is clear, and the recommendation fits the client. Everyone knows she has a good eye, but nobody has written down what she looks for. Asking AI to “review this” would automate the ambiguity.
On August 31, Microsoft described how one of its content leaders documented her review rubric and turned it into a repeatable pre-review inside the content workflow.
The benefit is not removing the expert. It is giving every draft the same basic inspection so the expert can spend time on judgment, coaching, and exceptions. The tradeoff is that a written rubric can freeze yesterday's preferences or reward tidy compliance over a genuinely strong idea.
One practical test: take ten recent examples the expert has already reviewed. Ask her to write five pass-or-revise questions, with one positive and one negative example for each. Let an assistant apply only those questions to five new drafts, then compare its flags with her independent review. Keep the rubric if it catches useful issues without narrowing the work; revise it when the disagreement teaches you something.
This is a vendor-authored account of Microsoft's own workflow, not an independent comparison, and its reported results do not predict savings in a smaller organization.
2. Put meeting help beside the conversation, not on top of it
A manager hears a number in a planning meeting and vaguely remembers a different figure from last quarter. Searching for it means dropping out of the discussion. Asking an AI aloud changes the rhythm for everyone. The useful assistant may be the one that offers a small, cited prompt at the edge of attention.
A paper submitted August 31 introduces InsightToast, a prototype that listens for information needs during data-rich meetings. It retrieves relevant material and presents short text or charts as temporary side-channel notifications.
The benefit is timely context without turning every uncertainty into a detour. The tradeoff is another system competing for attention, listening to the room, and deciding what deserves to appear. Even a correct prompt can steer a discussion simply because it arrived first.
A reasonable first move is much smaller than live transcription. For one recurring meeting, prepare three source-linked cards: the current target, the last decision, and one known uncertainty. Show a card only when the matching topic comes up. Afterward, ask whether each card resolved a question, distracted the group, or arrived too late. That test reveals the value of timely context before adding microphones or proactive retrieval.
The study used a small sample and one document-heavy scenario, and the paper presents a prototype rather than evidence of broad workplace gains. Privacy, consent, language, and accuracy also change the fit. If your team cannot explain what is heard, retained, and shown, a human-prepared side channel is the safer option.
3. Measure the context delivered, not the context promised
A developer sees that an agent supports a very large context window and assumes it will remember the instructions, tool results, and partial work from a long task. Halfway through, the agent repeats a search or loses a constraint. The invoice shows plenty of tokens; the result shows that the wrong material survived.
New research based on 55 archived coding-agent trajectories found that instructions, artifacts, tool output, and agent-created state behave differently in working memory. Tests of object-aware compression and retrieval found that gains seen during calibration did not always transfer to held-out tasks. The researchers also report that equal token budgets did not mean equal delivered context or equal memory-management cost.
The benefit of compression or retrieval is that a long-running agent can continue without carrying every old detail into every step. The tradeoff is hidden selection: saving tokens may remove the decision, failed attempt, or current artifact that the next action needs. A larger allowance can also increase cost without improving continuity.
One practical test: choose five completed agent tasks that required several tools or handoffs. Then note repeated work, lost constraints, correction time, and total usage. Compare outcomes before changing memory settings or buying a larger context tier.
The study is limited to coding agents and a modest trajectory set. It does not identify one memory strategy that will win across vendors or business workflows. It does offer a useful caution: a context-window specification describes capacity, not what the system selected, delivered, or used well.
The bigger pattern
Good assistance leaves the human work easier to see. A rubric exposes judgment, a side channel protects the meeting's flow, and a memory check shows which context actually reached the next step. Each adds some setup and none removes the need for an owner, but each can turn a vague promise into a practical test.
Which recurring task on your team depends most on one person's unwritten judgment, and what is the first review question you would write down?
Verified sources
- Microsoft Azure: Inside Microsoft's marketing team: Scaling expertise with AI
- Abolnejadian and Brehmer meeting-assistance research team: InsightToast: Proactive Information Retrieval & Glanceable Visualization in the Side Channel of Data-Rich Meetings
- Chen et al. coding-agent memory research team: Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents
Continue your decision path
Move from understanding to action.
AI Evaluation Rubric Builder
Turn this edition's decision points into a concrete working plan.
3 Things AI: The “Start With the Moment” Edition
Three fresh signals show when to stop an expensive test, why customer corrections belong in evaluation, and where urgent AI work should run.
Read nextApply this signal to your architecture.
Identify the workflow, context, and controls to structure first.
Open Architecture Assessment