Operating question
A smaller organization can gain more from checking the full chain, authority and final target than from giving an agent broader permission or choosing a model with a higher general score.
Agent Systems
Daily Signal: Let the workflow check itself before it acts
For
Leaders and workflow owners
You will leave with
4 operating decisions
Reading mode
11 min · 11 verified sources
Reading guide9 sections · Canadian briefing+
Highest-value moves
- 01Review the combined data flow and permissions of installed agent skills, not only each package in isolation.
- 02Separate prompts, rules, safety memory and tool permissions so a failure can produce one testable and reversible change.
- 03Put runnable acceptance tests and authority provenance in front of consequential work, then verify the final screen or external result.
- 04Choose model routes with bilingual task evidence for accepted outcomes, honest deferral, privacy, cost and staff effort.
Fresh research points Canadian SMEs toward chain-level skill review, versioned controls, executable tests and verified screen targets before AI changes shared work.
Today's strongest signal: the next useful AI gain is a workflow that checks its own proposed action before it touches shared work. Think of a familiar business situation. A service coordinator asks an assistant to read a request, choose the right procedure, update the customer record and send a confirmation. The answer can look polished while the real risk sits elsewhere: two harmless-looking tools combine into a risky chain, an approval points to the wrong authority, or the assistant clicks the button beside the one it intended.
Research released on August 10 makes that problem unusually concrete. One study shows that agent skills can appear safe alone and become harmful in combination. Another treats prompts, rules, memory and tool permissions as separate controls that can improve after observed failures. Security tests can help when shown before work begins, yet their coverage still matters. Screen-action research adds a simple idea: mark the proposed target, look again and then confirm. Safety measurements also fail when they score intent instead of the harmful outcome the business actually wants to stop.
For a Canadian small or medium-sized organization, the opportunity is practical. You may not need a larger model or a broader automation mandate. A useful first test may be a chain review, a versioned rule, an executable acceptance check or a second look before a consequential click. The tradeoff is more design and slower early runs. If a task is rare, easy to undo or already handled well by a person, the added control layer may cost more than it saves.
1. Review the installed chain, not only each tool
The fresh ColluSkill study tested six agent-skill scanners. It found a shared blind spot: several skills could pass individual inspection while their ordered combination created a harmful workflow. The reported attack succeeded 96.0% of the time on average against the tested scanners. A chain-aware defence reduced that rate to 22.5% while allowing 99.5% of benign workflows through. These are research results in controlled settings, not a guarantee for every commercial agent platform.
A skill is a packaged instruction or capability an agent can use. In a smaller firm, one skill might read files, another summarize them and a third send or update something. None looks dangerous in isolation. Risk appears when one leaves an artifact that another trusts, permissions accumulate, or a later step turns a draft into an external action.
Canada's toolkit for SME AI deployers is relevant because it distinguishes organizations integrating systems from those building the underlying model and recognizes their limited governance capacity. A useful first test is an installed-chain inventory. For one agent, list every skill, the data it reads, the artifact it creates, the next skill that can consume it and the strongest permission reachable across the whole chain. Rerun that review whenever a skill changes.
The opportunity is safe reuse of small capabilities. The tradeoff is added review work whenever the toolbox changes. A realistic reason not to install a third-party skill is that your team cannot inspect its instructions, dependencies or update path. Keep the capability manual until ownership and evidence are clear.
2. Turn failures into narrow control changes
SHE, a new safety-harness framework, separates an agent's system prompt, rule bank, safety memory and tool policy. It uses failed trajectories—records of what the agent saw and did—to identify which control needs adjustment. On the paper's benchmarks, the evolved harness cut the reported attack success rate by 3.1 times compared with a static harness and also transferred to held-out risks.
The practical point is not to let an agent rewrite its own rules in production. It is to stop treating every failure as a prompt-editing exercise. A rule about customer discounts belongs in a decision policy. A record of a suspicious pattern belongs in safety memory. Permission to send a payment belongs in tool policy. When those controls are separate, a team can change one, test it and roll it back without hiding the change inside a long prompt.
A useful first test is a weekly failure review for one bounded workflow. Save the input, proposed action, tool calls, approval result and external outcome. Classify the cause as missing instruction, wrong rule, stale memory, excess permission or poor execution. Propose the smallest change, test it against the failure and at least ten ordinary cases, then record who approved the new version.
The opportunity is learning from real operation without replacing the whole system. The tradeoff is maintenance: each control needs an owner and test set. If your team cannot retain safe traces or separate production evidence from sensitive customer data, do not build an automatic learning loop. Use a manual incident review with redacted examples.
3. Put acceptance tests in front of generation
A study of security tests as executable specifications examined 2,705 code-generation and repair trajectories across 31 tasks, 16 weakness categories and two model families. Showing all visible tests before generation improved hidden functional-and-security success by 19.3 percentage points on average. The benefit was not universal: two of nine benchmark-model conditions worsened, and candidates that passed every visible test still failed hidden behaviour families.
An executable specification is a check that can run and return a clear pass or fail. The idea applies beyond software. Before an assistant prepares a quote, tests can verify that the customer, currency, price list, approval level and expiry date are present. Before it updates inventory, tests can confirm a unique purchase order, expected item and allowed quantity range.
A useful first test is to write five acceptance checks before asking the assistant to produce anything. Include one ordinary case, one missing field, one conflicting record, one permission boundary and one outcome that must be rejected. Keep a second hidden set for evaluation so the assistant cannot merely shape its answer to the visible examples. Compare accepted results, false passes and rework with the current process.
The opportunity is clearer requirements and faster repair. The tradeoff is false confidence when the tests cover only the happy path. If the workflow depends mostly on judgment that cannot be expressed or reviewed consistently, automated tests may add ceremony without real protection. Keep the AI in a drafting role and invest first in a human review rubric.
4. Preserve authority as work moves between agents
The new POLIS multi-agent study ran 5,280 episodes across structured delegation and resource-allocation tasks. A provenance-aware guard—one that remembered where authority originated—recorded no violations in the paper's matched laundering cases. A guard that trusted only the currently visible local state admitted violations in 22 of 96 cases. The study also found that many workflows continued safely after a prohibited attempt was blocked.
Provenance means the origin and history of a piece of information or permission. This matters when an assistant summarizes an approval, another agent reformats it and a third agent acts. A sentence saying “approved” is not equivalent to a signed approval tied to a named decision, limit, version and expiry.
Ontario's Responsible Use of Artificial Intelligence Directive provides a useful provincial pattern: start with the problem, use proportionate controls, retain meaningful human oversight and give people a way to raise concerns. A useful first test is an authority envelope attached to every consequential handoff. Store the approver, exact scope, source record, time, expiry and permitted next action. If a transformation drops any field, the next agent can draft but cannot execute.
The opportunity is delegation without losing the decision boundary. The tradeoff is more structured data and occasional stops. A realistic reason not to use multiple agents is that one deterministic workflow with one approval can do the job more clearly. Extra agents are not a benefit when they only relay the same instruction.
5. Ask the interface agent to look again before it acts
LookAgain treats a predicted screen coordinate as a hypothesis. The system marks the proposed location, gathers a close visual view, and then confirms or revises the prediction. The researchers report improved results on general and refusal-aware interface benchmarks, especially where small targets, dense controls or unfamiliar layouts make a single prediction fragile.
This matters because many business automations operate through a visual interface when no stable application programming interface is available. A screen can move after a banner appears. Two customer rows may look similar. “Save draft” may sit beside “Send.” A correct plan can still cause the wrong action if the final coordinate is wrong.
Illustrative scenario: a 14-person property manager uses an assistant to prepare maintenance updates in a vendor portal. The assistant may open a work order and draft the status, but before clicking “notify tenant” it highlights the tenant name, unit, selected recipients and button label for confirmation. If the page moved or the fields do not match the work order, it stops without sending.
A useful first test is 50 replayable interface cases with moved controls, dense lists, modal windows, slow loading and a deliberately ambiguous target. Measure correct target selection, safe refusal, recovery time and unintended actions. The opportunity is help with systems that lack integrations. The tradeoff is slower execution and brittle visual evidence. If the action is financial, legal, safety-related or hard to reverse, prefer a supported API or keep the final action human.
6. Validate the control against the outcome it claims to prevent
A fresh audit of internal harmfulness scores distinguishes harmful intent in a prompt from a harmful outcome produced later by a specific model and decoding setup. In matched tests, wrapping increased harmful generation on one model while the prompt looked safer to the internal score. Among wrapped harmful prompts, the reported outcome ranking reversed, placing successful attacks below failed ones. The pattern persisted across three target models, seven attack families and two judges.
This is a measurement lesson. A dashboard can show a high safety score while the actual workflow still leaks data, sends an unauthorized message or changes the wrong record. A proxy is a measure used in place of the outcome we care about. Proxies are useful only when their relationship to that outcome is tested and keeps holding after the system changes.
Canadian privacy regulators' generative AI principles tell organizations to test necessity, proportionality, validity and reliability, keep accountability with the organization and provide meaningful human review for significant decisions. A useful first test is to name one prohibited outcome, such as disclosing a customer detail outside its approved purpose. Build representative attempts, run the complete deployed workflow and measure actual disclosure, not only a prompt-risk score.
The opportunity is a smaller set of measures that support real decisions. The tradeoff is harder evaluation and more careful data handling. If your team cannot observe the outcome safely, a polished proxy may be worse than an honest “not yet measured.” Keep the workflow bounded until evidence is possible.
7. Choose the route with a task scorecard, not one model ranking
A Dutch public-sector evaluation worked with domain experts to score more than 30 multilingual and Dutch-specific models across factuality, honesty, social bias, energy use, cost and training-data transparency. No model led across every dimension. The researchers also found that factuality and honesty were distinct: answering correctly did not imply acknowledging uncertainty well.
A second August 10 paper, Agentic Auto-Research is Fuzz Testing, offers a related distinction. Cheap signals can guide the next experiment, but protected final validation must decide whether the result is real. Guidance is not proof.
For a bilingual Canadian business, a useful first test is a task scorecard for one workflow in English and Canadian French. Include accepted-result accuracy, honest deferral, privacy boundary, cost per accepted result, response time and staff effort. Test two or three routes: perhaps a deterministic lookup, a smaller model and a stronger model. Route routine cases cheaply and reserve the expensive path for cases where it improves the accepted outcome.
The opportunity is lower cost and a better fit for each task. The tradeoff is ongoing evaluation as models and prices change. A realistic reason not to maintain several models is low volume: routing complexity can exceed the savings. In that case, choose one bounded route and keep a deterministic fallback.
Highest-value moves
- Inventory one agent's full skill chain and record the strongest permission reachable across it.
- Put five runnable acceptance checks and one authority envelope in front of a consequential workflow.
- Test the complete outcome—including the final screen or external result—before expanding automation.
Today's strongest thesis
A trustworthy AI workflow earns permission one checked handoff, one verified target and one measured outcome at a time.
Verified sources
- arXiv: ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners
- arXiv: SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
- arXiv: Security Tests as Executable Specifications for LLM Code Generation: Benefits, Trade-offs, and Coverage Limits
- arXiv: Multi-Agent AI Safety as an Institutional Design Problem
- arXiv: LookAgain: Closed-Loop GUI Grounding with Visually Grounded Reflection
- arXiv: Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks
- arXiv: From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
- arXiv: Agentic Auto-Research is Fuzz Testing
- Innovation, Science and Economic Development Canada: Toolkit for small- and medium-sized enterprises deploying artificial intelligence
- Office of the Privacy Commissioner of Canada: Principles for responsible, trustworthy and privacy-protective generative AI technologies
- Government of Ontario: Responsible Use of Artificial Intelligence Directive
Continue your decision path
Move from understanding to action.
Daily Signal: Build the rewind before the autopilot
Fresh evidence shows Canadian SMEs how checkpoints, structured handovers, abstention and grounded reports can make AI workflows easier to trust and recover.
Read nextApply this signal to your architecture.
Identify the workflow, context, and controls to structure first.
Open Architecture Assessment