Amazon Neptune now supports dual-stack mode, enabling database clusters to accept connections over IPv4, IPv6, or both protocols simultaneously. This allows organizations to adopt IPv6 while maintaining backward compatibility with existing IPv4 deployments.
Neptune dual-stack mode supports two configurations. Private dual-stack mode provides IPv6 endpoints that remain isolated from the internet, suitable for internal applications and private graph databases. Public dual-stack mode enables IPv6 endpoints accessible from the internet, supporting internet-facing applications and hybrid network environments. Clients connect seamlessly using their preferred protocol with no application changes required.
Dual-stack mode is available in all AWS Regions where Amazon Neptune is supported. To get started, see the Neptune setup documentation.
Amazon Neptune now supports dual-stack mode, enabling database clusters to accept connections over IPv4, IPv6, or both protocols simultaneously. This allows organizations to adopt IPv6 while maintaining backward compatibility with existing IPv4 deployments. Neptune dual-stack mode supports two configurations. Private dual-stack mode provides IPv6 endpoints that remain isolated from the internet, suitable for internal applications and private graph databases. Public dual-stack mode enables IPv6 endpoints accessible from the internet, supporting internet-facing applications and hybrid network environments. Clients connect seamlessly using their preferred protocol with no application changes required. Dual-stack mode is available in all AWS Regions where Amazon Neptune is supported. To get started, see the Neptune setup documentation.
AWS End User Messaging now supports rich media and interactive messaging for RCS across all 22 supported countries. With the new SendRcsMessage API, you can send rich cards, carousels, images, videos, and interactive suggestion buttons that let recipients take action directly inside their messaging app.
RCS message recipients can tap to confirm an appointment, browse a product catalog, complete a payment in a webview, share their location, or interact with an AI agent, all without leaving their phone’s messaging app. Behind each of these interactions is the same AWS infrastructure you already use to build applications. RCS becomes the interface layer that connects your backend services, your data, and your AI directly to your end users through your conversation with them.
With this release AWS now supports four RCS message types (text, files, rich cards, and carousels). These message types can be used with any combination of six actions (replies, URLs, webviews, phone calls, maps, and calendar events) to bring web and mobile app experiences directly into conversations.. Each message supports configurable SMS or MMS fallback for recipients without RCS.
AWS End User Messaging also introduces RCS Conversation pricing for 21 countries consisting of one flat rate for unlimited messages within a 24-hour session, so you can build back-and-forth workflows without per-message cost pressure.
RCS messaging is available in all AWS Regions where AWS End User Messaging is available. To learn more, see sending rich RCS messages in the AWS End User Messaging User Guide.
AWS End User Messaging now supports rich media and interactive messaging for RCS across all 22 supported countries. With the new SendRcsMessage API, you can send rich cards, carousels, images, videos, and interactive suggestion buttons that let recipients take action directly inside their messaging app.
RCS message recipients can tap to confirm an appointment, browse a product catalog, complete a payment in a webview, share their location, or interact with an AI agent, all without leaving their phone’s messaging app. Behind each of these interactions is the same AWS infrastructure you already use to build applications. RCS becomes the interface layer that connects your backend services, your data, and your AI directly to your end users through your conversation with them.
With this release AWS now supports four RCS message types (text, files, rich cards, and carousels). These message types can be used with any combination of six actions (replies, URLs, webviews, phone calls, maps, and calendar events) to bring web and mobile app experiences directly into conversations.. Each message supports configurable SMS or MMS fallback for recipients without RCS.
AWS End User Messaging also introduces RCS Conversation pricing for 21 countries consisting of one flat rate for unlimited messages within a 24-hour session, so you can build back-and-forth workflows without per-message cost pressure.
RCS messaging is available in all AWS Regions where AWS End User Messaging is available. To learn more, see sending rich RCS messages in the AWS End User Messaging User Guide.
Announcing Capability Insights for AWS, an open-source solution for regional capabilities
Today, AWS announces the launch of Capability Insights, an open-source solution that enables you to deploy regional capabilities data inside your own Amazon Virtual Private Cloud (VPC). This self-hosted dashboard addresses the needs of teams building multi-Region architectures requiring regional capabilities data deployed as infrastructure they own, inside their network, and under their governance. The solution is designed for organizations with data residency requirements, compliance teams needing internal reporting, and teams planning regional expansion or multi-Region recovery strategies.
The dashboard auto-refreshes every 24 hours with AWS capabilities data across all Regions, covering services, features, API operations, and CloudFormation resource types. The Workload Analysis component scans your AWS CloudTrail logs and AWS CloudFormation stacks to filter 200+ services down to the number of services your account actually uses, reducing multi-week gap analysis to quick reviews. All data remains within your VPC perimeter, supporting compliance and data residency requirements while providing full ownership and control over the infrastructure hosting the regional capabilities data.
Announcing Capability Insights for AWS, an open-source solution for regional capabilities
Today, AWS announces the launch of Capability Insights, an open-source solution that enables you to deploy regional capabilities data inside your own Amazon Virtual Private Cloud (VPC). This self-hosted dashboard addresses the needs of teams building multi-Region architectures requiring regional capabilities data deployed as infrastructure they own, inside their network, and under their governance. The solution is designed for organizations with data residency requirements, compliance teams needing internal reporting, and teams planning regional expansion or multi-Region recovery strategies.
The dashboard auto-refreshes every 24 hours with AWS capabilities data across all Regions, covering services, features, API operations, and CloudFormation resource types. The Workload Analysis component scans your AWS CloudTrail logs and AWS CloudFormation stacks to filter 200+ services down to the number of services your account actually uses, reducing multi-week gap analysis to quick reviews. All data remains within your VPC perimeter, supporting compliance and data residency requirements while providing full ownership and control over the infrastructure hosting the regional capabilities data.
Amazon Time Sync Service introduces support for microsecond accurate time on 26 additional EC2 instance types in all commercial regions. Built on Amazon’s proven network infrastructure and the AWS Nitro System, microsecond accurate time and nanosecond precision hardware timestamps leverage the reference clocks running in the Nitro System directly, enabling customers to easily order application events, measure 1-way network latency, and increase distributed application transaction speed.
Starting today, customers can access microsecond accurate time on these additional instance types by creating a Precision Time Placement Group (PTPG), a new placement strategy that allows customers to launch instances with Precision Time Protocol hardware clock (PHC) enabled. Customers that require both low network latency as well as precision time can associate a PTPG with their Cluster Placement Group (CPG), so that their low-latency workloads also benefit from microsecond accurate time.
Amazon Time Sync Service introduces support for microsecond accurate time on 26 additional EC2 instance types in all commercial regions. Built on Amazon’s proven network infrastructure and the AWS Nitro System, microsecond accurate time and nanosecond precision hardware timestamps leverage the reference clocks running in the Nitro System directly, enabling customers to easily order application events, measure 1-way network latency, and increase distributed application transaction speed. Starting today, customers can access microsecond accurate time on these additional instance types by creating a Precision Time Placement Group (PTPG), a new placement strategy that allows customers to launch instances with Precision Time Protocol hardware clock (PHC) enabled. Customers that require both low network latency as well as precision time can associate a PTPG with their Cluster Placement Group (CPG), so that their low-latency workloads also benefit from microsecond accurate time. For more information, refer to the Amazon Time Sync Service documentation.
Amazon WorkSpaces for agents is now generally available, enabling AI agents to securely access and operate desktop applications through managed WorkSpaces environments. Enterprises run critical business processes on desktop applications (ERP systems, CRMs, mainframes, and proprietary tools) where years of customization, undocumented logic, and strict compliance requirements make them too critical to abandon and costly to modernize. WorkSpaces for agentsnow gives AI agents a managed cloud workspace where they can see the screen and operate these applications the way humans do, without requiring application modernization or custom integrations.
WorkSpaces uses the same infrastructure for agents as organizations have trusted for over a decade to deliver secure, managed desktops at scale. Agents inherit the same identity controls, network isolation, and compliance boundaries as human users, so organizations gain automation without giving up governance. Organizations can automate workflows such as claims processing, patient record updates, trade settlement, and back-office operations. The service works with any agent framework using Model Context Protocol (MCP), and pricing scales based on active session time.
Since launching in Preview, customer and partner feedback has shaped new capabilities. MCP tool forwarding allows agents to interact with applications and the desktop operating system through direct MCP calls rather than using computer use tools, improving accuracy, reducing latency, and lowering cost. Real-time session control gives operators live visibility into agent activity with the ability to revoke access mid-session. Domain-joined fleet support lets agents operate under existing Active Directory identities, extending the same access policies and audit attribution that apply to employees.
Amazon WorkSpaces for agents is now generally available, enabling AI agents to securely access and operate desktop applications through managed WorkSpaces environments. Enterprises run critical business processes on desktop applications (ERP systems, CRMs, mainframes, and proprietary tools) where years of customization, undocumented logic, and strict compliance requirements make them too critical to abandon and costly to modernize. WorkSpaces for agentsnow gives AI agents a managed cloud workspace where they can see the screen and operate these applications the way humans do, without requiring application modernization or custom integrations.
WorkSpaces uses the same infrastructure for agents as organizations have trusted for over a decade to deliver secure, managed desktops at scale. Agents inherit the same identity controls, network isolation, and compliance boundaries as human users, so organizations gain automation without giving up governance. Organizations can automate workflows such as claims processing, patient record updates, trade settlement, and back-office operations. The service works with any agent framework using Model Context Protocol (MCP), and pricing scales based on active session time.
Since launching in Preview, customer and partner feedback has shaped new capabilities. MCP tool forwarding allows agents to interact with applications and the desktop operating system through direct MCP calls rather than using computer use tools, improving accuracy, reducing latency, and lowering cost. Real-time session control gives operators live visibility into agent activity with the ability to revoke access mid-session. Domain-joined fleet support lets agents operate under existing Active Directory identities, extending the same access policies and audit attribution that apply to employees.
To learn more, visit Amazon WorkSpaces for AI agents. To get started building, see the documentation and sample code on GitHub.
Escalar la disrupción del cibercrimen mediante la innovación y la IA
Por: Steven Masada, asesor jurídico adjunto, Unidad de Crímenes Digitales de Microsoft.
Microsoft ha adoptado un nuevo enfoque para combatir el ciberdelito, dirigiéndose a la cadena de suministro de ciberataques, no solo a los servicios individuales. En un caso que se ha desvelado de manera reciente, atacamos de manera simultánea dos herramientas de ciberdelincuencia utilizadas de manera amplia, Amadey y StealC, después de que un análisis asistido por IA revelara que dependen de la misma infraestructura.
Esta acción va tras la «cadena de montaje» del cibercrimen, donde herramientas coordinadas impulsan ransomware, fraudes financieros e interrupciones en los servicios públicos. Amadey y StealC se usan a menudo juntos: Amadey ayuda a los atacantes a acceder a dispositivos, mientras que StealC roba contraseñas e información sensible. Juntos, forman un eslabón crítico en la cadena. Solo en las dos primeras semanas de mayo, Amadey y StealC se vincularon a más de 140.000 ordenadores infectados en todo el mundo, lo que pone de manifiesto la amplitud con la que se utilizan.
En nuestro trabajo con Europol y socios industriales, nos dirigimos a ambas herramientas a la vez. El objetivo: romper la cadena. Desde el inicio de la operación, Microsoft ha identificado más de 18.000 ordenadores víctimas, ha cortado el control criminal de esos dispositivos y colabora con proveedores de telecomunicaciones para ayudar a proteger a los clientes afectados a nivel mundial.
Cuando varias partes de una operación se ven interrumpidas juntas, los ataques son más difíciles de lanzar, escalar y recuperarse. El resultado: menos servicios interrumpidos, menos oportunidades para que los ciberdelincuentes se beneficien y más fricciones cuando intentan reconstruirse.
Ya no basta con ir tras las amenazas una a una. Tenemos que interrumpir cómo se organizan los ataques.
Lo nuevo es cómo combinamos el análisis de IA con un uso ampliado de esa ley.
Amadey y StealC fueron desarrollados por ciberdelincuentes independientes, pero dependían de la misma infraestructura. Para entender cómo funcionaban, los investigadores utilizaron IA, incluido Copilot, para analizar con rapidez el malware, a través de hacer preguntas en inglés sencillo en lugar de revisar de manera manual código complejo. Eso ayudó a sacar a la luz detalles clave, descubrir datos ocultos y probar resultados en una fracción del tiempo, para convertir lo que habría llevado horas o días en minutos y permitir al equipo detectar conexiones más rápido.
Esas conclusiones permitieron al equipo legal tratar ambas familias de malware como parte de una sola conspiración. En lugar de ir tras cada herramienta por separado, como hemos hecho en el pasado, usamos RICO para acusar a múltiples facilitadores cómplices implicados en toda la operación. En total, la Unidad de Crímenes Digitales de Microsoft interrumpió más de 200 servidores de mando y control, los sistemas que los delincuentes utilizan para controlar dispositivos infectados, robar datos y mantener los ataques en marcha.
Al agrupar las herramientas, podemos interrumpir la cadena del cibercrimen de manera más eficiente y eficaz, de una forma que refleje mejor cómo funcionan en realidad estas redes hoy en día.
El cibercrimen ahora funciona como una cadena de montaje
El cibercrimen ya no es una serie de ataques aislados, es un sistema coordinado.
Herramientas especializadas gestionan cada paso: uno obtiene acceso, otro roba credenciales y otros venden o explotan ese acceso para fraude, ransomware, espionaje u otros fines nefastos. Diferentes actores pueden estar involucrados en cada etapa, pero juntos convierten el acceso en beneficio, con rapidez y a gran escala.
Cómo las herramientas de ciberdelincuencia están diseñadas para ser modulares
Esa estructura también crea un punto de vulnerabilidad. Las personas detrás de estas herramientas cibercriminales quizá nunca interactúen de manera directa, pero sus herramientas están diseñadas para funcionar juntas. Si se pueden identificar esas conexiones, se pueden interrumpir múltiples fases de un ataque a la vez.
Cómo se desarrollan estos ataques en el mundo real
La mayoría de la gente nunca oirá los nombres Amadey o StealC, pero sienten los efectos. Un hospital bloqueado de sistemas críticos. Una ciudad incapaz de ofrecer servicios esenciales. Una pequeña empresa que pierde acceso a las cuentas de la noche a la mañana. Un jubilado que perdió todos sus ahorros.
Estos ataques no ocurren todos de golpe. Se desarrollan paso a paso: los atacantes entran, roban contraseñas, el acceso se reutiliza o vende, y a veces se reutiliza para operaciones más específicas. Por ejemplo, Microsoft ha observado al actor afiliado a Rusia Secret Blizzard aprovechar las infecciones de Amadey para desplegar malware personalizado contra objetivos en Ucrania.
Al atacar varios puntos de esa cadena a la vez, reducimos la posibilidad de que un solo compromiso se convierta en un daño generalizado. En resumen: menos ataques tienen éxito y menos personas sienten el impacto cuando lo hacen.
Ninguna organización puede hacer esto sola
Acciones como esta subrayan una realidad fundamental: tenemos éxito cuando colaboramos. Ninguna organización, ya sea gubernamental o industrial, tiene visibilidad completa sobre cómo operan las amenazas cibernéticas a través de fronteras y sectores. Lo que hace que este esfuerzo sea efectivo es la combinación de perspectivas y datos.
Reunir esos esfuerzos amplió nuestros conjuntos de datos colectivos y permitió identificar las conexiones entre ambas herramientas y actuar con rapidez sobre ellas. Ese entendimiento compartido permitió una respuesta coordinada que fue más allá de lo que cualquier organización individual podría lograr por sí sola.
Esto demuestra por qué las asociaciones importan. La industria comparte conocimientos técnicos, el gobierno aporta visibilidad y necesitamos formas fiables de intercambiar esa información. Solo si se trabaja desde la misma perspectiva podremos mantenernos por delante de los atacantes, para interrumpir no solo herramientas individuales sino también los sistemas que hacen posible el cibercrimen.
Crear una presión sostenida sobre la ciberdelincuencia
Este trabajo no termina con una sola acción. Los ciberdelincuentes se adaptan con rapidez, por eso mantenemos el seguimiento de la evolución de estas operaciones y colaboramos con socios para interrumpirlas.
La interrupción autorizada por el tribunal de Microsoft en este caso se combina con los esfuerzos continuos para rastrear cómo los ciberdelincuentes reconstruyen, identificar nuevas infraestructuras y trabajar con socios para interrumpir los servicios de los que dependen para operar. También incluye incorporar los hallazgos de esta interrupción en iniciativas como el programa Statutory Automated Disruption de Microsoft, que ayuda a acelerar la eliminación de dominios e infraestructuras maliciosas.
El objetivo no es solo detener una operación, sino ralentizar el propio sistema, para hacer que los ataques sean más difíciles de lanzar, escalar y recuperarse. Al combinar la visión impulsada por IA, acciones legales y sólidas alianzas, podemos seguir con el aumento del coste del cibercrimen y reducir su impacto.
Durante más de una década, la Unidad de Delitos Digitales (DCU) de Microsoft ha trabajado para combatir el cibercrimen y las amenazas de los estados-nación, ha presentado alrededor de 40 casos desde 2008 y colaborado con las fuerzas del orden para desmantelar redes criminales. Descubran más sobre los esfuerzos del equipo aquí.
Amazon SageMaker Inference now supports container image caching, enabling up to 2x faster end-to-end scaling for generative AI models during scale-out events. When your endpoint scales out, the service pre-caches your container image so new instances can start serving traffic faster, without waiting for large container images to be pulled from Amazon ECR.
Generative AI workloads typically use large container images (10 GB or more) for deep learning frameworks and model serving. Previously, every new instance launched during scale-out had to pull the full image from ECR, adding several minutes of cold-start latency. Container image caching eliminates this bottleneck by pre-pulling the image so new instances launch with the container already available locally. Customers don’t need to make any changes. The service automatically caches whatever image URI is specified in your endpoint or inference component configuration. This capability supports accelerator instance types, single-model endpoints, and inference component-based endpoints.
With this launch, SageMaker Inference now offers a comprehensive scaling optimization suite for generative AI: sub-minute concurrency metrics for up to 6x faster load detection, instance-store container caching for faster scaling on existing instances, and container image caching for up to 2x faster scaling on new instances.
Container image caching is available in all AWS commercial regions where SageMaker Inference is supported. To learn more, visit the launch blog.
Amazon SageMaker Inference now supports container image caching, enabling up to 2x faster end-to-end scaling for generative AI models during scale-out events. When your endpoint scales out, the service pre-caches your container image so new instances can start serving traffic faster, without waiting for large container images to be pulled from Amazon ECR. Generative AI workloads typically use large container images (10 GB or more) for deep learning frameworks and model serving. Previously, every new instance launched during scale-out had to pull the full image from ECR, adding several minutes of cold-start latency. Container image caching eliminates this bottleneck by pre-pulling the image so new instances launch with the container already available locally. Customers don’t need to make any changes. The service automatically caches whatever image URI is specified in your endpoint or inference component configuration. This capability supports accelerator instance types, single-model endpoints, and inference component-based endpoints. With this launch, SageMaker Inference now offers a comprehensive scaling optimization suite for generative AI: sub-minute concurrency metrics for up to 6x faster load detection, instance-store container caching for faster scaling on existing instances, and container image caching for up to 2x faster scaling on new instances. Container image caching is available in all AWS commercial regions where SageMaker Inference is supported. To learn more, visit the launch blog.
IAM Identity Center now enables customer managed applications to programmatically access AWS accounts on behalf of their users, including the ability to discover accounts and roles assigned to a user and retrieve temporary credentials required for AWS account access.
If you have a customer managed application that authenticates users through an external identity provider (IdP), you can configure that IdP as a trusted token issuer (TTI) in IAM Identity Center. With this launch, you can now enable AWS account access for this application. Users who have already signed in through the IdP can access their assigned AWS accounts and obtain temporary security credentials for their authorized roles without a separate authentication flow. This eliminates redundant sign-in prompts that previously required users to re-authenticate even after signing in through their external identity provider.
This feature is available for organization instances of IAM Identity Center. IAM Identity Center administrators must explicitly enable AWS account access for each customer managed application. Only management account administrators or delegated administrators can enable this capability, ensuring centralized governance over which applications can access account-level resources.
This feature is available in all commercial AWS Regions, the AWS GovCloud (US) Regions, and the China Regions. To get started, navigate to the IAM Identity Center console, select your customer managed application, and enable AWS account access. For more information, see Enable AWS account access for customer managed applications in the IAM Identity Center User Guide.
IAM Identity Center now enables customer managed applications to programmatically access AWS accounts on behalf of their users, including the ability to discover accounts and roles assigned to a user and retrieve temporary credentials required for AWS account access.
If you have a customer managed application that authenticates users through an external identity provider (IdP), you can configure that IdP as a trusted token issuer (TTI) in IAM Identity Center. With this launch, you can now enable AWS account access for this application. Users who have already signed in through the IdP can access their assigned AWS accounts and obtain temporary security credentials for their authorized roles without a separate authentication flow. This eliminates redundant sign-in prompts that previously required users to re-authenticate even after signing in through their external identity provider.
This feature is available for organization instances of IAM Identity Center. IAM Identity Center administrators must explicitly enable AWS account access for each customer managed application. Only management account administrators or delegated administrators can enable this capability, ensuring centralized governance over which applications can access account-level resources.
This feature is available in all commercial AWS Regions, the AWS GovCloud (US) Regions, and the China Regions. To get started, navigate to the IAM Identity Center console, select your customer managed application, and enable AWS account access. For more information, see Enable AWS account access for customer managed applications in the IAM Identity Center User Guide.
AWS GovCloud (US) now offers Claude Opus 4.8 — Anthropic’s most capable generally available model to date — delivering meaningful advances across agentic coding, professional knowledge work, and long-running autonomous tasks for developers and enterprises building production AI applications.
Claude Opus 4.8 can perform longer autonomous runs, deeper reasoning, and consistency to be trusted with production work. For coding, the Opus 4.8 reads codebases like an engineer, plans before it edits, and holds context across long sessions in real repositories. For agentic tasks, it is better at finding paths around obstacles instead of stalling, recovering from its own errors, and knowing when to ask for help versus when to keep going. For knowledge work, it better synthesizes across long documents and complex sources, self-checks its output, and delivers structured deliverables that hold up to review.
Amazon Bedrock keeps your data within AWS infrastructure and provides access to Claude Opus 4.8 through a unified service with AWS-managed features like Guardrails, Knowledge Bases, and regional data residency. To learn more, see Amazon Bedrock documentation and regional availability.
AWS GovCloud (US) now offers Claude Opus 4.8 — Anthropic’s most capable generally available model to date — delivering meaningful advances across agentic coding, professional knowledge work, and long-running autonomous tasks for developers and enterprises building production AI applications.
Claude Opus 4.8 can perform longer autonomous runs, deeper reasoning, and consistency to be trusted with production work. For coding, the Opus 4.8 reads codebases like an engineer, plans before it edits, and holds context across long sessions in real repositories. For agentic tasks, it is better at finding paths around obstacles instead of stalling, recovering from its own errors, and knowing when to ask for help versus when to keep going. For knowledge work, it better synthesizes across long documents and complex sources, self-checks its output, and delivers structured deliverables that hold up to review.
Amazon Bedrock keeps your data within AWS infrastructure and provides access to Claude Opus 4.8 through a unified service with AWS-managed features like Guardrails, Knowledge Bases, and regional data residency. To learn more, see Amazon Bedrock documentation and regional availability.
Two new models are now available in the Kiro IDE and CLI for the AWS GovCloud (US-West) Region.
OpenAI GPT-5.4 is now available in Kiro for complex reasoning, coding, document analysis, and multi-step agentic workflows. It helps developers build AI applications and production workflows that can interpret context, interact with tools, operate software environments, and verify outputs across multiple steps. GPT-5.4 runs on Amazon Bedrock’s next-generation inference engine with isolated queues and durable execution for resilient workloads. Available with a 272K context window and 1.2x credit multiplier.
NVIDIA Nemotron 3 Super 120B is now available in Kiro as an open weight model option. A hybrid mixture-of-experts model activating only 12B of its 120B parameters for high compute efficiency and fast inference on agentic tasks. 256K context window with 32K max output. Available with a 0.25x credit multiplier.
Ensure your IDE or CLI is updated to the latest version, then restart it to access the new models from the model selector. For more details about Kiro in AWS GovCloud (US), visit the GovCloud documentation or contact your AWS account team for more information. To learn more about Kiro, visit the Kiro product page.
Two new models are now available in the Kiro IDE and CLI for the AWS GovCloud (US-West) Region. OpenAI GPT-5.4 is now available in Kiro for complex reasoning, coding, document analysis, and multi-step agentic workflows. It helps developers build AI applications and production workflows that can interpret context, interact with tools, operate software environments, and verify outputs across multiple steps. GPT-5.4 runs on Amazon Bedrock’s next-generation inference engine with isolated queues and durable execution for resilient workloads. Available with a 272K context window and 1.2x credit multiplier. NVIDIA Nemotron 3 Super 120B is now available in Kiro as an open weight model option. A hybrid mixture-of-experts model activating only 12B of its 120B parameters for high compute efficiency and fast inference on agentic tasks. 256K context window with 32K max output. Available with a 0.25x credit multiplier. Ensure your IDE or CLI is updated to the latest version, then restart it to access the new models from the model selector. For more details about Kiro in AWS GovCloud (US), visit the GovCloud documentation or contact your AWS account team for more information. To learn more about Kiro, visit the Kiro product page.