AI NewsAI NewsAuto PublishingAWS AgentCoreRuntime instancesGPU持久会话AWSAgentCoreAI Agents

AWS AgentCore Adds Persistent Runtime Instances for Production Agents

Managed EC2 infrastructure, GPU support, and sessions of up to 14 days target long-running agent work.

ENHE AI5 min0 views
AWS AgentCore Adds Persistent Runtime Instances for Production Agents

Key takeaways

AWS announced AgentCore Runtime instances on August 6, 2026. The feature provides persistent, managed EC2 infrastructure for production AI agents, with multi-agent collaboration, GPU support, and sessions lasting up to 14 days. That addresses long-running state and resource continuity, but it does not remove operational responsibility. A safe first trial asks whether a task truly needs hours or days of state, then uses minimal permissions, non-sensitive data, an automatic termination rule, and a cost record covering CPU, GPU, idle time, network access, and session duration. Teams should validate isolation, logging, human approval, backup, and rollback before connecting a persistent runtime to real production data.

AWS announced persistent Runtime instances on August 6.
They support multi-agent work, GPUs, and sessions up to 14 days.
Persistence does not remove governance.
Trials should measure cost, access, and termination.

# AWS AgentCore Adds Persistent Runtime Instances for Production Agents

August 9, 2026

On this page

  • Direct answer
  • Fact sources
  • Action guide
  • Why it matters
  • Impact
  • FAQ
  • Sources

Direct answer

Try persistent Runtime instances only when a task needs long-lived state or multi-agent collaboration. Begin with non-sensitive data and minimal permissions, and measure session duration, idle cost, GPU use, network access, and termination behavior.

Fact sources

AWS announced AgentCore Runtime instances on August 6, 2026.

The announcement describes managed persistent EC2 infrastructure, multi-agent collaboration, GPU support, and sessions up to 14 days.

Long-lived compute increases the importance of cost, data, network, cleanup, and rollback controls.

Five steps for an AgentCore persistent-runtime trial

  1. Confirm that the workload needs cross-hour or cross-day state.
  2. Use a least-privilege role, non-sensitive data, and automatic termination.
  3. Record session duration, CPU/GPU, idle time, network access, and cost.
  4. Test state isolation, logs, and human approval for multi-agent work.
  5. Expand only after budget, data boundaries, and rollback are clear.

Why it matters

Agent deployment is shifting from individual model calls to persistent runtimes. Longer sessions enable complex work while making resource leaks, permission drift, and cost surprises more consequential.

Impact for ordinary AI users

A 14-day capability does not mean every task should run continuously. Enterprises should treat termination, backups, logs, and data cleanup as part of the workload design.

Related tools and tutorials

ENHE software and local-deployment pages support runtime choices; account services and skill-learning pages support permission, budget, and rollback checks.

AI software and tool entry points · AI account permissions and cost services · AI skill tutorials and validation methods · AI frontier news overview

FAQ

Is 14 days guaranteed everywhere?

The announcement describes a capability; verify region, account, quota, and version availability.

Are persistent instances always cheaper?

No. Idle time and GPU consumption can increase cost.

Can I connect a production database first?

Start with isolated data and least privilege, then expand after evidence.

Source links

  • AWS News Blog: Runtime instances for AgentCore (2026-08-06)
  • AWS Docs: AgentCore runtime instances
  • AWS Docs: AgentCore runtime tools

What this means for everyday users

Put session duration, idle resources, GPUs, network access, cost, and termination ownership on the launch checklist.

Related reading

Cloudflare Unifies Workers AI and AI Gateway: Choosing an AI Control Plane

Cloudflare announced on August 7, 2026 that Workers AI and AI Gateway are moving toward one AI control plane. The announcement describes unified bindings, observability, billing, and dynamic routing across Cloudflare-managed GPUs and external providers. That can simplify multi-model operations, but it does not guarantee lower cost, consistent quality, or compliance. Choose the control plane around a real task: define model and latency needs, data sensitivity, budget, routing and fallback requirements, then test one low-risk endpoint with read-only logs. Preserve the actual model and version, latency, error, cost, permission, and fallback evidence before moving broader traffic. Small applications may be better served by one provider and clear logs until routing complexity has a measurable benefit.

GitHub Copilot Usage Metrics Adds Agent-App Activity

GitHub announced on August 7, 2026 that the Copilot Usage Metrics API now reports activity from third-party agent apps. Enterprise, organization, enterprise-user, and organization-user reports can expose the activity in one-day and 28-day windows. The new totals_by_3rd_party_agent data includes a stable agent_id and a display name that may change; the identifier should be the join key. This gives administrators a finer view of cost, permissions, and workflow adoption, but it does not automatically explain business value. Start with a read-only sample, reconcile time zones, pagination, and overlapping windows, then associate agent activity with AI credits, members, repositories, and permission changes before changing budgets or access.

AWS Shows AgentCore Policy Workflows with Tenant Isolation and Versioned Skills

AWS’s August 7, 2026 machine-learning case study describes how Cohere Health uses Amazon Bedrock AgentCore to turn clinical prior-authorization policies into structured data. The architecture combines Runtime microVM isolation, Gateway for unified tool access, Memory for session history, and the Agent Skills open standard for versioned domain capabilities. Skills are evaluated with reference data and expert review before release. The reusable lesson is not to automate medical judgment with one prompt. It is to separate tenants, tools, data sources, versions, feedback, and approval, then begin with public or de-identified documents before connecting sensitive business data. Keep the same evidence trail when the workflow changes.

How to Audit AI Discoverability with Cloudflare Agent Readiness and AEO

Cloudflare announced Agent Readiness and Answer Engine Optimization tools on August 6, 2026. Agent Readiness checks whether agents can discover, read, and call a site, while AEO measures whether assistants recommend or cite it for realistic category questions. This guide turns the announcement into six repeatable checks: inspect robots and sitemaps, publish machine-readable facts and sources, document APIs or agent interfaces, run unbranded customer prompts, record citations and competitor mentions, and change one variable at a time. The metrics are diagnostic samples, not search rankings or guaranteed market share; preserve model, prompt, date, and page-version evidence. Repeat the scan after each material content or access change.

Cloudflare Previews WebMCP: Give Browser Agents Site Tools

Cloudflare announced a WebMCP developer preview on August 6, 2026. A site can enable tool packs in the Cloudflare Dashboard so browser AI agents can discover and call actions through a standard surface instead of guessing buttons and parsing human-oriented HTML. The preview injects a bridge at the edge, runs tools in the visitor’s browser, and can reuse the visitor’s existing session for a site MCP endpoint. Because it is a preview, users should start with a test account, minimal tool packs, non-critical actions, and explicit confirmation before allowing messages, purchases, or account changes. Recheck permissions whenever the browser or pack version changes.

Google Expands Gemini API Managed Agents with 3.6 Flash and Hooks

Google’s July 28, 2026 announcement expands Gemini API Managed Agents with Gemini 3.6 Flash, Hooks, and additional trigger capabilities. Google positions the service as a way to build more reliable, production-ready agents, but managed infrastructure does not remove the need for evaluation, permissions, logging, or cost controls. A practical first trial fixes the model version and region, enables only the tools the task needs, and uses a read-only or reversible workflow. Record trigger behavior, retries, latency, token use, failures, and human approvals before allowing external messages, database writes, or expensive calls. Re-run the same test after every model or trigger change.

Summary

AgentCore persistent Runtime instances bring long-running agents closer to production infrastructure, while amplifying cost and access responsibility. Validate real long-state needs before scaling.

Sources

Table of contents

Latest Insights