The agent's instructions now learn from its own mistakes. amazon bedrock's agentcore prompt…
the agent's instructions now learn from its own mistakes. amazon bedrock's agentcore prompt optimizer, out september 16, mines production traces for a reward signal and auto-rewrites the system prompt, with every proposal gated by platform guardrails.
prompt engineering just became a feedback loop.
Context
An AWS Machine Learning blog of 16 September 2026, a technical companion to a launch post, says the system prompt optimizer in Amazon Bedrock AgentCore uses agent traces recorded in AgentCore Observability together with a reward signal to produce an improved system prompt. A reflector reads evaluated traces and returns proposed edits to the agent configuration, and before any proposal can be applied it must pass platform-level guardrails that screen for drift such as prompts growing longer, quoting trace text verbatim or relaxing safety constraints. The workflow is propose, validate by offline batch evaluation and online A/B testing, then promote. The AgentCore docs say you specify a target evaluator as the reward signal and that recommendations are generated by LLMs and should be reviewed and tested before applying.
Traces, a reward signal and guardrails are supported. Auto-rewrites the system prompt is narrower in the sources: the service produces a recommended prompt with an explanation, AWS says to review and test it, and promotion goes through evaluation and A/B testing, so it is not a silent rewrite. The blog is dated 16 September but the feature's launch date was not in the text read, so out September 16 is unverified as a release date. The reported results, 95.83 percent on AppWorld and 79.15 percent on WebShop against GEPA and MIPROv2, are vendor-reported and belong to the experimental open-source Sub-Agent Reflector, while the feature available today uses a Single Agent Reflector whose scores were not read. That prompt engineering became a feedback loop is the author's opinion.
Related work
- Earlier note on AgentCore evals in GitHub Actions ↗The evaluation side that supplies the reward signal.
Watch next
- The dated launch post and independent results for the Single Agent Reflector.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 19:16 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →