Skip to main content

AI & assistant-friendly summary

This section provides structured content for AI assistants and search engines. You can cite or summarize it when referencing this page.

Summary

Harness GA June 17, Agents Classic cutoff July 30, Payments GA August 18, Memory FGAC August 28. Harness vs Runtime, Gateway Policy, and when transacting agents still should not own checkout.

Key Facts

  • •Harness GA June 17, Agents Classic cutoff July 30, Payments GA August 18, Memory FGAC August 28
  • •AWS lifecycle notice (June 30, 2026) — Amazon Bedrock Agents Classic is now Bedrock Agents Classic, in maintenance for new customers after July 30, 2026
  • •Net-new agent builds should use Bedrock AgentCore
  • •On June 17, 2026, AgentCore Harness reached general availability — a managed, config-driven agent loop on the same platform as Runtime, Memory, Gateway, and Identity (What's New)
  • •First-party signals we reuse here — Gateway canary on a B2B CRM assistant (12 tools, ~8k turns/day): median tool round-trip ~180 ms → ~95 ms after server-side Gateway (Gateway post)

Entity Definitions

Amazon Bedrock
Amazon Bedrock is an AWS service discussed in this article.
Bedrock
Bedrock is an AWS service discussed in this article.
Lambda
Lambda is an AWS service discussed in this article.
DynamoDB
DynamoDB is an AWS service discussed in this article.
CloudWatch
CloudWatch is an AWS service discussed in this article.
IAM
IAM is an AWS service discussed in this article.
API Gateway
API Gateway is an AWS service discussed in this article.
foundation model
foundation model is a cloud computing concept discussed in this article.

Amazon Bedrock AgentCore: The Production Guide for Net-New AI Agents on AWS

AI AgentsPalaniappan P9 min read

Quick summary: Harness GA June 17, Agents Classic cutoff July 30, Payments GA August 18, Memory FGAC August 28. Harness vs Runtime, Gateway Policy, and when transacting agents still should not own checkout.

Key Takeaways

  • Harness GA June 17, Agents Classic cutoff July 30, Payments GA August 18, Memory FGAC August 28
  • AWS lifecycle notice (June 30, 2026) — Amazon Bedrock Agents Classic is now Bedrock Agents Classic, in maintenance for new customers after July 30, 2026
  • Net-new agent builds should use Bedrock AgentCore
  • On June 17, 2026, AgentCore Harness reached general availability — a managed, config-driven agent loop on the same platform as Runtime, Memory, Gateway, and Identity (What's New)
  • First-party signals we reuse here — Gateway canary on a B2B CRM assistant (12 tools, ~8k turns/day): median tool round-trip ~180 ms → ~95 ms after server-side Gateway (Gateway post)
Amazon Bedrock AgentCore: The Production Guide for Net-New AI Agents on AWS
Table of Contents

AWS lifecycle notice (June 30, 2026) — Amazon Bedrock Agents Classic is now Bedrock Agents Classic, in maintenance for new customers after July 30, 2026. Net-new agent builds should use Bedrock AgentCore. Full matrix: lifecycle roundup.

On June 17, 2026, AgentCore Harness reached general availability — a managed, config-driven agent loop on the same platform as Runtime, Memory, Gateway, and Identity (What’s New). Combined with the June 30 lifecycle batch that puts Agents Classic into maintenance for new customers after July 30, 2026, the default path for net-new production agents on AWS is no longer Classic action groups.

This guide is the architecture map for that path: what AgentCore is in late 2026, when to pick Harness vs Runtime, how Memory / Gateway / Identity / Payments fit together, and what an honest Classic cutover looks like. Deep dives on pricing components, Gateway server-side tools, and AgentCore vs Quick Suite TCO stay in their own posts — link out, do not duplicate.

First-party signals we reuse here — Gateway canary on a B2B CRM assistant (12 tools, ~8k turns/day): median tool round-trip ~180 ms → ~95 ms after server-side Gateway (Gateway post). First-party TCO silhouette (July 2026): ~500-employee Quick Suite stack ~$3,580/mo vs AgentCore support-style agent at 50K sessions/mo ~$791/mo platform + Haiku-like inference (decision guide).

Reproduce this — Model platform spend on the AgentCore pricing calculator (list rates as of 2026-07-04, us-east-1). For Gateway cutover shape, use the public artifacts: server-side vs client matrix and Responses Gateway sketch.

What AgentCore Is Now (Not a Classic Wrapper)

Amazon Bedrock AgentCore is a modular agent platform: build, deploy, and operate agents with any supported framework and foundation model, with managed isolation, memory, tools, identity, and observability (AWS overview).

ServiceRole
HarnessConfig-driven managed loop — model, prompt, tools, memory; microVM sessions with filesystem/shell
RuntimeServerless host for custom agent code (Strands, LangGraph, CrewAI, LlamaIndex, OpenAI Agents SDK, BYO container); MCP/A2A
MemoryShort-term + long-term memory with extractable strategies; FGAC via Gateway Cedar (Aug 28, 2026); flexible namespace keys
GatewayAPIs, Lambda, and MCP servers as tools; OAuth; Policy interception; customer-configurable rate limits
IdentityWorkload + user identity; Cognito / Okta / Entra ID; credential brokering
BrowserManaged browser for web interaction
Code InterpreterSandboxed code execution
ObservabilityOTEL-compatible traces into CloudWatch
PolicyCedar (or NL→Cedar) gates on Gateway tool calls
EvaluationsQuality scoring on sessions/traces — batch, online, and A/B testing GA (June 2026); skill evaluators (August 2026)
PaymentsOpt-in microtransactions for paid APIs/MCP/content — GA August 18, 2026; not a checkout replacement

Mentioned, not covered deep here: Registry (org catalog, including cross-account RAM sharing), Optimization (eval-driven config experiments), Coinbase Marketplace billing for wallet usage. See What this post doesn’t cover.

Opinionated take: treat AgentCore as the paved road for customer-facing and product-embedded agents. Treat Amazon Quick Suite as the paved road for employee permission-aware knowledge work. Hybrid is normal at 200+ employees — not indecision.

Two Entry Paths: Harness vs Runtime

                    ┌─────────────────┐
                    │  Your product   │
                    │  / API layer    │
                    └────────┬────────┘
                             │
              ┌──────────────┴──────────────┐
              ▼                             ▼
     ┌────────────────┐            ┌────────────────┐
     │ AgentCore      │            │ AgentCore      │
     │ Harness        │            │ Runtime        │
     │ (config loop)  │            │ (your code)    │
     └────────┬───────┘            └────────┬───────┘
              │                             │
              └──────────────┬──────────────┘
                             ▼
         Memory · Gateway · Identity · Browser · CI · Observability · Payments

Choose Harness when:

  • You want a production agent from configuration in hours, not a framework repo
  • Tools are Gateway/MCP/inline and the loop is standard reason→act→observe
  • You may later export to Strands on the same platform without re-platforming memory/tools

Choose Runtime when:

  • You already ship LangGraph / CrewAI / Strands / custom Python and need AWS isolation + identity
  • You need long-running async agents, multi-agent A2A, or a custom container image
  • Orchestration logic is the product (complex branching, human-in-the-loop graphs)

What broke — Early 2026 prototype path that treated AgentCore as “enable memory on a Classic agent alias.” The team kept InvokeAgent + Classic action groups, then discovered Gateway Policy, Identity JWT inbound, and Harness/Runtime were the documented production surfaces — not Classic aliases. Detection: architecture review against the AgentCore developer guide before a compliance audit. Fix: rebuild tools as Gateway targets and pick Runtime (existing Python agent) instead of forcing Classic. Lesson: lifecycle banners are not a migration recipe; Classic and AgentCore are different control planes.

Memory: Strategies, Not Just a DynamoDB Flag

AgentCore Memory supports short-term (session) and long-term (cross-session) stores. Long-term extraction is driven by strategies — built-in (semantic, summarization, user preference, episodic), overrides, or self-managed pipelines (memory strategies).

Harness can provision managed memory with defaults (semantic + summarization, event expiry) so you are not inventing a summary table on day one. Runtime frameworks integrate via the Memory APIs and SDKs.

Design rules that matter in production:

  1. Namespace memory by user (or tenant) — never a shared global notebook for multi-tenant SaaS. As of August 28, 2026, AgentCore Memory supports flexible namespace variables (up to five custom keys such as organization, tenant, team, or environment) so you do not duplicate strategies per tenant.
  2. Enforce isolation at the infrastructure layer. Memory FGAC fronts Memory with a Gateway OAuth target and Cedar policies so each caller only hits their own actorId or namespace derived from JWT claims — do not keep that check in agent code. Batch Memory APIs are not covered by FGAC; do not use them as the multi-tenant write path.
  3. Set retention / TTL early — Memory event and retrieval charges compound when facts live forever (pricing guide).
  4. Keep Knowledge Bases for documents; keep Memory for interaction state. Duplicating PDFs into Memory is a cost and quality bug. CreateEvent now also accepts a json payload type (up to 100 KB) for non-conversational behavioral events — still not a document store.

Tools: Gateway, MCP, and Policy

Gateway turns OpenAPI, Smithy, Lambda, and existing MCP servers into agent-callable tools with ingress/egress auth. For Responses API server-side execution (discover → select → execute without a client tool loop), see the Gateway server-side tools post — including the ~180 → ~95 ms median tool RTT benchmark on a 12-tool CRM assistant.

Policy sits on Gateway: Cedar (or natural language converted to Cedar) intercepts tool calls before execution — who can call which tool with which arguments. That is the control you want for write tools in finance and healthcare, not a prompt that says “please be careful.” Gateway also supports customer-configurable rate limits scoped to JWT claims, IAM principals, targets, tools, or models — use them for multi-tenant fairness instead of an application-layer token bucket.

Opinionated take: once you cross ~10 tools or need shared OAuth SaaS connectors, prefer Gateway over ad-hoc Lambda wiring in the app. Keep a client confirmation loop only when product requires human approval before every write.

Identity and Session Isolation

Every Runtime/Harness session runs in an isolated microVM. Session A must not leak conversation state or filesystem into Session B.

AgentCore Identity ties the agent workload identity to enterprise IdPs (Cognito, Okta, Entra ID, Auth0). For per-user credential scoping into third-party tools (Google, Slack, etc.), prefer inbound JWT/OAuth on the harness/runtime so Identity can bind tokens to the end user — SigV4 caller paths do not get the same on-behalf-of story today (Harness security).

Scope IAM execution roles per agent. An over-privileged agent role is automated lateral movement with a chat UI.

Observability and Quality

Observability emits OpenTelemetry-compatible telemetry into CloudWatch — step-level visibility across model calls, tools, and memory operations. Forward to Datadog/New Relic/etc. only if that stack speaks OTEL and you accept dual-ingest cost.

Traces tell you what happened. Evaluations tell you whether it was good. Wire evals into CI before you scale session volume. Batch evaluations, online evaluations, and A/B testing are generally available as of June 2026; August 2026 added skill evaluators (Builtin.SkillSelectionAccuracy, Builtin.SkillInstructionFollowing) and managed DeepEval/AutoEval judges. Confirm regional availability before committing a launch gate — do not treat traces alone as a quality number.

Production Reference Architecture

A durable customer-facing assistant on AgentCore typically looks like this:

  1. Edge — API Gateway / ALB + your auth (or JWT passed through to AgentCore Identity).
  2. Entry — Harness (InvokeHarness) or Runtime (InvokeAgentRuntime) — not Classic InvokeAgent for net-new.
  3. Memory — user-scoped STM/LTM with TTL and strategies matching the product (preferences vs episodic).
  4. Gateway — tool catalog (Lambda/OpenAPI/MCP) + Policy for write paths.
  5. Optional — Browser / Code Interpreter only on turns that need them (always-on burns duration).
  6. Observability + Evaluations — CloudWatch + nightly/weekly eval suites (GA — not a preview checkbox).
  7. Payments (opt-in) — only if the agent must settle pay-per-use APIs/MCP/content; session budget + Cedar, never unbounded checkout.
  8. Model — Bedrock (or OpenAI-compatible / Gemini via Harness provider support) chosen for price-performance; switchable without rewriting tool IAM if Gateway owns tools.
# Conceptual control-plane shape (not a copy-paste SDK pin).
# Confirm field names against the AgentCore API reference for your SDK version.
#
# Path A — Harness: CreateHarness / UpdateHarness / InvokeHarness
#   model + system prompt + tool/Gateway refs + memory config
#
# Path B — Runtime: CreateAgentRuntime + container/framework artifact
#   JWT authorizer (optional) + execution role + VPC config
#
# Shared: Gateway targets, Memory resource, Identity credential providers
# Optional: Payment Manager + session budget (not a checkout replacement)

Payments: GA for Agent-to-API Spend, Not Checkout

Previous approach: transacting agents were a preview control plane (May 2026) and this guide treated Payments as out of scope. Teams either skipped paid APIs or wired a wallet in application code.

What AWS changed: on August 18, 2026, AgentCore Payments reached general availability. It orchestrates x402 (including the upto scheme) and the Machine Payments Protocol (MPP), integrates Coinbase and Stripe Privy wallets, and enforces maxSpendAmount + expiry on each PaymentSession at the infrastructure layer (Payments developer guide). Policy (who may call which paid tool) and payment sessions (how much may be spent) are orthogonal levers.

What this enables: research or browser agents that buy a paywalled API or MCP call inside a budgeted session, with CloudWatch vended logs/spans on every data-plane payment.

When it matters: you already have a Gateway tool catalog and need pay-per-use endpoints without standing up a merchant-of-record wallet.

When it does not: card-on-file checkout, refunds, or any flow where the shopper’s payment instrument must stay on your existing processor. Do not put PAN in Memory. Do not treat Payments GA as permission to let a shopping agent complete an unbounded purchase. eCommerce copilot still assembles a cart and hands off to checkout — see when agents choose products.

Cost Reality (Platform vs Model)

Do not budget AgentCore as “DynamoDB GB-month plus per-invocation.” As of the rates wired into our calculator (2026-07-04, us-east-1):

ShapeHow it meters (list)
Runtime / Browser / Code InterpreterActive vCPU-hour + GB-hour
MemorySTM events, LTM events/month, retrievals (per 1k)
GatewayInvokes / search / tools indexed

Model tokens remain the largest line for most agents. Lean Runtime+Memory deployments often land roughly 1.15× model spend; full Browser+CI+Gateway+Identity stacks closer to 1.4× — directional, not a guarantee. Work the numbers on the AgentCore pricing calculator and the 12-components post.

Classic → AgentCore: Honest Cutover Posture

WorkloadPosture
Net-new agentBuild on Harness or Runtime
Existing Classic in productionKeep running; no forced rewrite on July 30
New customers after July 30, 2026Plan AgentCore; Classic is maintenance for new customers
“Flip the alias in a day”Reject — inventory tools, Gateway/Policy, Memory, Identity first

Classic tutorials on this site (agentic workflows, Classic tool-use, Classic multi-agent supervisor) remain useful for understanding the old loop. They are not the recommended implementation path for net-new builds. Prefer Gateway + Harness/Runtime; use the multi-agent supervisor post as pattern language, not as Classic CloudFormation you copy into production in Q3 2026.

What to Do This Week

  1. Decide Harness vs Runtime with one sentence of justification (config vs custom code).
  2. Inventory tools; if ≥10 or multi-team, draft Gateway targets and Policy for writes.
  3. Define Memory namespaces + TTL before the first production user. If the agent is multi-tenant, front Memory with Gateway FGAC instead of an in-agent authz check.
  4. Run the pricing calculator with your sessions × active seconds × peak memory.
  5. If the agent must pay for tools, set a PaymentSession budget and Cedar deny on ProcessPayment for any role that can raise that budget. If it must not, leave Payments unconfigured.
  6. If you still have a Classic prototype, schedule a cutover design review — not a silent alias rename.

What This Post Doesn’t Cover

  • Exact unit prices beyond calculator-as-of dates — always check AWS AgentCore pricing.
  • Identity federation deep-dive (token vault, custom claims, multi-IdP) — follow-on post.
  • Evaluations CI design and Optimization A/B experiment design — follow-on (the products themselves are GA).
  • Payments connector setup (Coinbase Marketplace subscription, Stripe Privy credentials, x402 upto vs MPP choice) — follow-on; this post only covers when to opt in.
  • Registry (including cross-account RAM sharing) and Optimization product setup.
  • EU AI Act classification for autonomous write tools.
  • GovCloud / every region matrix — confirm regional availability in AWS docs before procurement language.

We have not re-benchmarked Harness cold-start p95 across every commercial region on the same day as this update — measure in your target Region before committing UX copy that promises “instant” first messages.


Need a production readiness assessment (Harness vs Runtime, Gateway Policy, Memory TTL, cost model)? Contact FactualMinds — AWS Select Tier Partner with AgentCore delivery experience across financial services, healthcare, and SaaS product agents.

Frequently asked questions

Is AgentCore just a wrapper around Bedrock Agents Classic?
No. That framing is outdated. AgentCore is a modular agent platform — Harness (config-driven managed loop), Runtime (serverless host for Strands, LangGraph, CrewAI, and custom containers), plus Memory, Gateway, Identity, Browser, Code Interpreter, Observability, Policy, and Evaluations. Agents Classic (the November 2023 orchestration API) enters maintenance for new customers after July 30, 2026. Net-new builds should start on AgentCore, not on Classic action groups plus an "AgentCore alias."
Should I use AgentCore Harness or AgentCore Runtime?
Use Harness when you want a production agent from configuration — model, system prompt, tools, memory — without owning the orchestration loop, and you are fine exporting to Strands later if you outgrow config. Use Runtime when you already have (or need) custom agent code on Strands, LangGraph, CrewAI, LlamaIndex, or a BYO container, and you need microVM session isolation, long-running async agents, or multi-agent protocols (MCP/A2A). Do not pick Runtime "for flexibility" if Harness covers the workflow — you inherit framework upgrade and loop-debugging cost.
What is the pricing model for AgentCore?
AgentCore is consumption-based. Runtime, Browser, and Code Interpreter bill primarily on active compute (vCPU-hour and GB-hour for session time). Memory bills on short-term and long-term events plus retrievals. Gateway bills on invocations, semantic search queries, and indexed tools. Identity, Observability, Evaluations, and Payments are separate opt-in lines. Model tokens are always a separate Bedrock line and usually dwarf platform spend. As of July 2026 us-east-1 list rates in our calculator: $0.0895/vCPU-hour and $0.00945/GB-hour for compute-shaped lines. Model your workload on the AgentCore pricing calculator and read the 12-components pricing post — do not assume DynamoDB GB-month as the Memory unit.
How does AgentCore Memory differ from Bedrock Knowledge Bases?
Complementary systems. Knowledge Bases is semantic search over documents (embeddings, retrieval at query time). AgentCore Memory is stateful conversation and user context — short-term within a session and long-term across sessions with extractable strategies (semantic, summarization, user preference, episodic). Think Knowledge Bases as the reference library and Memory as the notebook. Most production agents need both.
Can I migrate an existing Agents Classic agent to AgentCore in a day or two?
Usually no for anything beyond a thin demo. Classic action groups, prompt templates, and InvokeAgent session attributes do not map 1:1 onto Harness config or a Runtime container. Plan a cutover: inventory tools → Gateway targets or MCP servers → choose Harness vs Runtime → re-validate IAM and Policy → canary traffic. Existing Classic deployments keep running; the July 30, 2026 change is about net-new customers, not forced teardown. A one-day flip is a red flag that you skipped Gateway Policy and Memory strategy design.
When should I NOT use AgentCore?
Skip AgentCore for single-turn Bedrock Converse/InvokeModel calls, read-only RAG that Knowledge Bases already solves, or employee-only knowledge work where Amazon Quick Suite wins on connectors and seats (see the AgentCore vs Quick decision guide). Also avoid if you need hard sub-second first-token SLAs without warm-session design — microVM cold starts are measured in seconds, not milliseconds.
What observability does AgentCore provide?
AgentCore Observability emits OpenTelemetry-compatible telemetry into Amazon CloudWatch so you can inspect execution paths, tool calls, and failures in one place. You can forward OTEL-compatible signals to third-party stacks that ingest that format. Build dashboards and alarms on latency, errors, and tool failure rates. Pair with AgentCore Evaluations (batch evaluations, online evaluations, and A/B testing are generally available as of June 2026; skill evaluators and third-party DeepEval/AutoEval judges shipped in August 2026) for quality scoring — do not treat traces alone as a quality gate.
Can I run AgentCore in a VPC?
Yes. Harness and Runtime sessions can connect to your VPC for private access to internal APIs, databases, and endpoints. Combine with Gateway Policy (Cedar) so tool calls are gated before they reach private resources. Treat Browser and Code Interpreter as high-risk surfaces — restrict networks and prompts even inside a VPC.
Should production agents use AgentCore Payments now that it is GA?
Only when the agent must pay for APIs, MCP servers, or content — not as a replacement for your existing checkout. AgentCore Payments reached general availability on August 18, 2026, with Coinbase and Stripe Privy wallets, infrastructure spend limits per PaymentSession, and x402/MPP protocol orchestration. Keep card capture and unbounded purchase on the merchant checkout. Use Payments for gated microtransactions with a session budget; skip it for WISMO, catalog, and support agents that never move money.

Reference bounds

How the AgentCore production guide stays bounded

Level 2 — reference architecture. Risk high. Oversight: approval required.

Separate traces, identity, gateway, and evaluation for a net-new agent on AgentCore.

Explore, then act, then confirm, then verify. Explore gathers the context for the AgentCore production guide. Act stays reversible. Confirm stops before an irreversible step. Verify is a separate check of the outcome.

Starts when
A team is standing up AgentCore for a new agent.
Tools
A write goes through a router, a permission check, a policy check, and a budget check. The model does not commit it. An approval token lives in tool context, not in the user message.
Stops for a person
Managed primitives do not replace the application policy or the human gate.
Checked by
Observability is on, and an evaluation suite exists, before session volume grows.
Untrusted data
Samples that ship with a long-lived key in the prompt.
If it fails
If the AgentCore production guide stops, name the reason: completed, budget_exceeded, timed_out, cancelled, guardrail_blocked, approval_required, tool_failure, verification_failed, or partial_completion. Retry a timeout at most twice. Back off on rate limits. After repeated verification failure, escalate. Stop when the budget is exhausted or a permission is denied. No unbounded loop. Managed primitives do not replace the application policy or the human gate.

This is the reference architecture for the page, not a published production deployment. The shared contract is the AWS store-agent architecture. Permissions and data boundaries are in securing store agents.

Palaniappan P
Palaniappan P

AWS Cloud Architect & AI Expert

AWS-certified cloud architect and AI expert with deep expertise in cloud migrations, cost optimization, and generative AI on AWS.

AWS ArchitectureCloud MigrationGenAI on AWSCost OptimizationDevOps

Recommended Reading

Explore All Articles »