Skip to main content

AI & assistant-friendly summary

This section provides structured content for AI assistants and search engines. You can cite or summarize it when referencing this page.

Summary

On Sep 18, 2026 AgentCore microVM Runtime V2 went GA — AWS P75 cold starts 1.9–2.0s vs 5.4–30s on V1, but consumption rates rose to $0.1276/vCPU-hour and $0.0169/GB-hour. Memory reclaim waits 120 seconds.

Key Facts

  • •On Sep 18, 2026 AgentCore microVM Runtime V2 went GA — AWS P75 cold starts 1
  • •9–2
  • •0s vs 5
  • •4–30s on V1, but consumption rates rose to $0
  • •1276/vCPU-hour and $0

Entity Definitions

Amazon Bedrock
Amazon Bedrock is an AWS service discussed in this article.
Bedrock
Bedrock is an AWS service discussed in this article.
EC2
EC2 is an AWS service discussed in this article.
serverless
serverless is a cloud computing concept discussed in this article.
IaC
IaC is a cloud computing concept discussed in this article.
CloudFormation
CloudFormation is a development tool discussed in this article.
CDK
CDK is a development tool discussed in this article.

AgentCore Runtime V2 GA: 1.9s P75 Cold Starts — Unit Rates Went Up

AI AgentsPalaniappan P7 min read

Quick summary: On Sep 18, 2026 AgentCore microVM Runtime V2 went GA — AWS P75 cold starts 1.9–2.0s vs 5.4–30s on V1, but consumption rates rose to $0.1276/vCPU-hour and $0.0169/GB-hour. Memory reclaim waits 120 seconds.

Key Takeaways

  • On Sep 18, 2026 AgentCore microVM Runtime V2 went GA — AWS P75 cold starts 1
  • 9–2
  • 0s vs 5
  • 4–30s on V1, but consumption rates rose to $0
  • 1276/vCPU-hour and $0
Night server aisle: a near rack of compact nodes lit in amber-gold, the far aisle receding into navy shadow
Table of Contents

On September 18, 2026, AWS announced general availability of the next AgentCore Runtime on serverless microVM compute. You opt in per runtime with platformVersion: V2. V1 stays the default if you omit the field on create.

This is not Runtime Instances (EC2 capacity providers, 14-day sessions, GA August 6, 2026). V2 is a start-and-bill change on the microVM you already run: snapshot restores instead of full image start, and unused memory reclaimed after 120 seconds instead of held until the session ends.

Opinionated take: Flip V2 for cold-start consistency — AWS published P75 1.9–2.0 seconds on 200 MB–2 GB images versus 5.4–30 seconds on V1. Do not flip because the What’s New page said “lower costs.” Consumption list rates went up (CPU $0.0895 → $0.1276/vCPU-hour, memory $0.00945 → $0.0169/GB-hour). Reclaim only helps sessions that actually idle ≥ 120 seconds after a memory spike. Short support turns never hit that window.

Platform map, Memory/Gateway/Identity, and Classic cutover stay in the AgentCore production guide. Component lines stay in the 12-components pricing post.

First-party signals we reuse here — Gateway canary on a B2B CRM assistant (12 tools, ~8k turns/day): median tool round-trip ~180 ms → ~95 ms after server-side Gateway (Gateway post). First-party TCO silhouette (July 2026): AgentCore support-style agent at 50K sessions/mo ~$791/mo platform + Haiku-like inference (decision guide). Those agents are microVM-shaped and short-session; V2 does not change Gateway or seat math, and a 3–5 second turn never reaches 120-second memory reclaim.

Reproduce this — Score v1-vs-v2-decision-matrix.md. Fill v1-peak-vs-v2-elastic-cost-worksheet.csv (AWS published V2 example plus two labeled models). Run v2-canary-checklist.md. The AgentCore pricing calculator still uses V1 peak-memory rates dated 2026-07-04 — do not treat it as a V2 modeler.


What AWS shipped on September 18, 2026

DimensionmicroVM V1 (default)microVM V2 (platformVersion)
Cold startFull start per instance; AWS P75 5.4–30s from 200 MB–2 GB imagesRestore a prepared snapshot; AWS P75 1.9–2.0s across that image range
Memory billingUnused memory held until the session ends (peak-shaped)Starts small; grows on demand; unused reclaimed after 120s; 128 MB floor
CPU billingActive vCPU-hour; scales to zero on I/O waitSame CPU-to-zero behavior; higher consumption rate
us-east-1 consumption$0.0895/vCPU-hour, $0.00945/GB-hour$0.1276/vCPU-hour, $0.0169/GB-hour
Create / updateREADY in secondsSnapshot prep: minutes; poll to READY / FAILED
Health checkStandard pingFirst healthy /ping within 120s or create fails; that ping is the snapshot
Env vars4 KB1.5 KB direct code, 2.5 KB container (limit to be raised to match V1)
IaCExisting CFN/CDK pathsCloudFormation and CDK cannot set platformVersion yet

Regions (V2): us-east-1, us-east-2, us-west-2, eu-west-1, ap-northeast-1. Docs: platform versions. Pricing: AgentCore pricing. Snapshot code rules: optimize V2.

A committed baseline at $0.0997/vCPU-hour and $0.0132/GB-hour is listed as launching by October 2026. It is not applied in the worksheet.


The cost trap: higher unit rates, reclaim on a 120-second clock

AWS’s own V2 pricing example is a customer-support agent: 1 million sessions/month, 10 minutes wall-clock, 90% I/O wait, 1 vCPU while active, memory walking 1 GB → 2 GB → 2.5 GB. Published V2 compute: $0.006703/session, $6,703/month.

That is not automatically cheaper than V1. V2 CPU is 42% higher per active vCPU-hour; V2 memory is 79% higher per GB-hour. Break-even on memory requires billed average GB around 56% of the V1 peak (0.00945 / 0.0169). The 10-minute AWS example still bills enough GB-seconds that modeled V1 peak-held memory at old rates undercuts published V2 (~$5,429 vs $6,703). The What’s New claim of “lower costs” only holds when reclaim actually drops average GB far below peak.

SilhouetteV1 Runtime computeV2 Runtime computeWhat it teaches
AWS published 1M × 10 min support exampleModeled ~$5,429/mo (peak 2.5 GB × 10 min at V1 rates)$6,703/mo (AWS published)10-minute sessions: rate hike can beat reclaim
50K × 5s chat turns (existing TCO shape)Modeled $6.87/mo RuntimeModeled $10.03/mo RuntimeSub-120s turns: +46%, reclaim never fires
120 × 8h sessions, 2.5 GB spike then 128 MB floorModeled $31.27/moModeled $14.89/moLong idle-after-spike: V2 wins ~52% on Runtime

The 50K / 8h rows are modeled list-rate math, not client bills. Open the CSV for the assumptions. Browser and Code Interpreter still list V1-shaped CPU/memory rates on the same pricing page — do not copy V2 Runtime rates onto those lines.

Choose V2 when:

  • Image size or concurrency makes V1 start time a tail (AWS 5.4–30s vs V2 ~2s).
  • Sessions idle ≥ 120 seconds after a memory spike (human wait, long tool/LLM I/O).
  • You have scored the matrix ≥ 14 and will canary with Cost Explorer tags.

Stay on V1 when:

  • Median turns are seconds (the 50K-session support shape).
  • Starts are already fine and you have not modeled the rate hike.
  • Primary Region is outside the five-Region V2 list.

Snapshot correctness — what breaks if you treat V2 like V1

V2 restores a snapshot taken after the first healthy /ping. Work at import / process start is copied onto every instance and frozen until the next create/update. Work in the /invocations handler runs per request.

What broke — Design-review failure, not a published client outage: a Gateway tool catalog (or STS session, or uuid/random seed, or time.monotonic() start mark) loaded at module scope “because it is slow.” After V2 restore, every instance shares the same catalog version, the same request-id entropy, or a clock that does not advance across restore. AWS is explicit: hostname is localhost and PID is 1 on every restored instance, so using either as a worker id collapses logs and locks. Detection: identical request IDs across sessions, expired credentials on a “fresh” instance, tool list stuck at snapshot time. Fix: refresh credentials and IDs in the handler; do not cache Gateway inventories at startup; BYO containers need snapshot-safe OpenSSL (openssl-snapsafe-libs on Amazon Linux 2023). Direct-code deployments already ship a snapsafe base image.

Sockets opened at startup do not survive restore; client caches (endpoint resolution, pools) do. Expect the first call after restore to reconnect. Do not bind a fixed source port.


Enable V2 without pretending IaC supports it

Assumes AWS CLI v2 and bedrock-agentcore-control in a V2 Region. Replace the role ARN and ECR URI. Docs: Enable V2.

# AWS CLI — create a microVM runtime on platform version V2
aws bedrock-agentcore-control create-agent-runtime \
  --agent-runtime-name "my-agent" \
  --role-arn "arn:aws:iam::111122223333:role/AgentExecutionRole" \
  --agent-runtime-artifact '{"containerConfiguration":{"containerUri":"111122223333.dkr.ecr.us-east-1.amazonaws.com/my-agent:latest"}}' \
  --network-configuration '{"networkMode":"PUBLIC"}' \
  --platform-version V2

The create call returns while status is still CREATING. Poll get-agent-runtime until READY or *FAILED. A second update before that returns ConflictException. Confirm with --query platformVersion.

# boto3 bedrock-agentcore-control — create on V2, then poll.
# platformVersion is not in the create response; call get_agent_runtime.
import time
import boto3

client = boto3.client("bedrock-agentcore-control", region_name="us-east-1")

created = client.create_agent_runtime(
    agentRuntimeName="my-agent",
    roleArn="arn:aws:iam::111122223333:role/AgentExecutionRole",
    agentRuntimeArtifact={
        "containerConfiguration": {
            "containerUri": "111122223333.dkr.ecr.us-east-1.amazonaws.com/my-agent:latest"
        }
    },
    networkConfiguration={"networkMode": "PUBLIC"},
    platformVersion="V2",
)

runtime_id = created["agentRuntimeId"]
while True:
    got = client.get_agent_runtime(agentRuntimeId=runtime_id)
    status = got["status"]
    if status == "READY" or status.endswith("FAILED"):
        break
    time.sleep(5)

assert got.get("platformVersion") == "V2"

Omit platformVersion on create → V1. Omit it on update → keep the current platform version. CloudFormation and CDK cannot set the field yet — Console, CLI, or SDK only until AWS adds it.


Canary without rewriting InvokeAgentRuntime

  1. Baseline (Day 0) — session-duration histogram, vCPU-hour / GB-hour, P75 start, tool errors, $/successful session on V1.
  2. Score — decision matrix; proceed if sum ≥ 14 or a hard cold-start tail.
  3. Model — worksheet. If median session is under 120s and starts are already fine, stop.
  4. Canary (72 hours) — one non-prod runtime tagged MigrationWave=agentcore-runtime-v2-canary; keep a V1 control; include a deploy that rebuilds the snapshot (checklist).
  5. Promote or park — promote only runtimes that share the image-size / idle-after-spike shape. Keep V1 for bursty sub-120s chat.

What This Post Doesn’t Cover

  • Runtime Instances (EC2 capacity providers, 14-day sessions) — Instances GA post.
  • Full AgentCore platform map — production guide.
  • Live V2 rates in the public calculator — still V1 peak-memory as of 2026-07-04 at the AgentCore pricing calculator.
  • Committed-baseline GA — listed as launching by October 2026; not modeled as live.
  • A re-run of AWS’s 1.9–2.0s P75 bench — we cite the What’s New numbers; measure yours on the canary.
  • A first-party V2 production engagement — there are still zero published AI-agent case studies. Proof here is the artifacts plus AWS published figures.

What to Do This Week

  1. Inventory runtimes by median session length versus 120 seconds, container image size, and primary Region.
  2. Score each candidate on the decision matrix — canary only if sum ≥ 14 or cold starts already hurt.
  3. Model V1 peak-held versus V2 elastic on the worksheet before flipping production.
  4. Audit startup code against the optimize-V2 rules (handler-side IDs, clocks, credentials; snapsafe OpenSSL in BYO images).
  5. Canary one non-prod runtime for 72 hours with the checklist. Promote only if $/successful session or P75 start wins.

Need help choosing V1 vs V2 vs Runtime Instances for a production agent fleet? Contact FactualMinds — AWS Select Tier Partner with Bedrock AgentCore delivery across SaaS and regulated workloads — or start from generative AI on AWS.


Frequently asked questions

What is AgentCore Runtime V2?
Announced generally available on September 18, 2026, Runtime V2 is a platformVersion on the existing serverless microVM compute in Amazon Bedrock AgentCore — not a new product and not Runtime Instances (EC2 capacity providers). V2 starts each instance from a prepared snapshot and reclaims unused session memory after 120 seconds of idle. V1 remains the default if you omit platformVersion on create.
When should I NOT flip platformVersion to V2?
Do not flip if (1) median sessions are a few seconds and never idle 120 seconds — reclaim never fires and list rates are higher, (2) your Region is outside us-east-1, us-east-2, us-west-2, eu-west-1, and ap-northeast-1, (3) you have not modeled V2 $0.1276/vCPU-hour and $0.0169/GB-hour against V1, or (4) startup still caches short-lived credentials, Gateway tool catalogs, or random seeds. Short chat agents on V2 are a rate hike, not a discount.
Are V2 unit rates lower than V1?
No. Consumption list rates in us-east-1 went up: CPU $0.0895 to $0.1276 per vCPU-hour, memory $0.00945 to $0.0169 per GB-hour. AWS can still show a lower bill when billed average memory is well below peak because unused memory is reclaimed after 120 seconds. A committed-baseline discount ($0.0997 / $0.0132) is listed as launching by October 2026 — it is not live in this post. Our AgentCore pricing calculator still uses V1 peak-memory rates dated 2026-07-04; use the worksheet in this post for V2 math.
How fast are V2 cold starts versus V1?
AWS reported a P75 cold start of 1.9 to 2.0 seconds for container images from 200 MB to 2 GB, compared with 5.4–30 seconds on V1. V2 prepares the environment once and restores the snapshot instead of repeating full startup. We did not re-run that bench; treat those numbers as AWS published, then measure P75 in your account on the canary.
What could go wrong after adopting V2?
Create/update can sit in CREATING/UPDATING for minutes and fail if /ping is not healthy within 120 seconds. Values computed at startup are frozen in the snapshot — expired STS, identical request IDs, and a Gateway tool catalog that never refreshes. Environment variables over 1.5 KB (direct code) or 2.5 KB (container) fail ValidationException (V1 allowed 4 KB). CloudFormation and CDK cannot set platformVersion yet, so an IaC-only pipeline silently stays on V1.
Is Runtime V2 the same as Runtime Instances?
No. V2 is a platform version on serverless microVM Runtime. Runtime Instances (GA August 6, 2026) run agents on EC2 types you pick via a capacity provider, with sessions up to 14 days and a separate EC2-plus-management bill. Stay on microVM V1 or V2 for bursty short sessions; use Instances when you need GPU or wall-clock beyond the 8-hour microVM design point.
How do I enable V2 on an existing runtime?
Set platformVersion to V2 on create-agent-runtime or update-agent-runtime (Console, CLI, or SDK). If you omit the field on create, the runtime uses V1. If you omit it on update, the runtime keeps its current platform version. Poll get-agent-runtime until status is READY or ends in FAILED, then confirm platformVersion is V2. Do not issue another update while the runtime is CREATING or UPDATING — that returns ConflictException.
Does V2 change session isolation or the 8-hour microVM limit?
No. Hardware-enforced session isolation, scale-to-zero, and the 8-hour microVM session design point remain. CPU still scales to zero during I/O wait. What changed is how instances start (snapshot restore) and how unused memory is billed (reclaim after 120 seconds, 128 MB floor).

Conceptual bounds

How the Runtime V2 stays bounded

Level 1 — conceptual. Risk medium. Oversight: monitored.

Account for Runtime V2 cold start and unit rates when you budget an agent.

Starts when
A team is planning Runtime V2.
Tools
This page does not grant tools. A later build still needs a named catalog, and anything unlisted stays unreachable.
Stops for a person
A faster start does not remove step limits or the write gate.
Checked by
Cost is modeled from the published rates and your session shape, not from a universal saving.
Untrusted data
Release notes treated as proof the agent is safe.
If it fails
Retry a timeout at most twice. Back off on rate limits. After repeated verification failure, escalate. Stop when the budget is exhausted or a permission is denied. No unbounded loop.

This page describes a bound. It is not a production agent and not a published deployment. The shared contract is the AWS store-agent architecture. Permissions and data boundaries are in securing store agents.

Palaniappan P
Palaniappan P

AWS Cloud Architect & AI Expert

AWS-certified cloud architect and AI expert with deep expertise in cloud migrations, cost optimization, and generative AI on AWS.

AWS ArchitectureCloud MigrationGenAI on AWSCost OptimizationDevOps

Recommended Reading

Explore All Articles »