Publicado el Deja un comentario

Stability AI Image Services now available in Amazon Bedrock

Amazon Bedrock announces the availability of Stability AI Image Services, a comprehensive suite of 9 specialized image editing tools designed to accelerate professional creative workflows. Stability AI Image Services enable granular control over image editing with a range of tools designed to work with your creative process, allowing you to take a single concept from ideation to finished product with precision and flexibility.

Stability AI Image Services offers two categories of image editing capabilities: Edit tools: Remove Background, Erase Object, Search and Replace, Search and Recolor, and Inpaint let you make targeted modifications to specific parts of your images. Control tools: Structure, Sketch, Style Guide, and Style Transfer give you powerful ways to generate variations based on existing images or sketches.

Stability AI Image Services is now available in Amazon Bedrock through the API and is supported in US West (Oregon), US East (N. Virginia), and US East (Ohio). For more information on supported regions, visit the Amazon Bedrock Model Support by Regions guide. For more details about Stability AI Image Services and its capabilities, visit the Stability AI product page and Stability AI documentation page

 

​Amazon Bedrock announces the availability of Stability AI Image Services, a comprehensive suite of 9 specialized image editing tools designed to accelerate professional creative workflows. Stability AI Image Services enable granular control over image editing with a range of tools designed to work with your creative process, allowing you to take a single concept from ideation to finished product with precision and flexibility. Stability AI Image Services offers two categories of image editing capabilities: Edit tools: Remove Background, Erase Object, Search and Replace, Search and Recolor, and Inpaint let you make targeted modifications to specific parts of your images. Control tools: Structure, Sketch, Style Guide, and Style Transfer give you powerful ways to generate variations based on existing images or sketches. Stability AI Image Services is now available in Amazon Bedrock through the API and is supported in US West (Oregon), US East (N. Virginia), and US East (Ohio). For more information on supported regions, visit the Amazon Bedrock Model Support by Regions guide. For more details about Stability AI Image Services and its capabilities, visit the Stability AI product page and Stability AI documentation page.   

Publicado el Deja un comentario

OpenAI open weight models expand to new regions on AWS Bedrock

Today, AWS announces the expansion of OpenAI open weight models on AWS Bedrock to eight new regions. This expansion brings these powerful AI models closer to customers in various parts of the world, enabling lower latency and improved performance for a wide range of AI-powered applications.

With this expansion, the OpenAI open weight models are now available in the following AWS Regions: US East (N. Virginia), Asia Pacific (Tokyo), Europe (Stockholm), Asia Pacific (Mumbai), Europe (Ireland), South America (São Paulo), Europe (London), and Europe (Milan), in addition to the previously supported region of US West (Oregon). This broader availability allows more customers to leverage these state-of-the-art AI models while keeping their data within their preferred geographic locations, helping to address data residency requirements and reduce network latency.

To learn more about OpenAI open weight models on AWS Bedrock and how to get started, visit the Amazon Bedrock console or check out our documentation. For more information about the initial release of these models on AWS Bedrock, refer to our previous blog post

 

​Today, AWS announces the expansion of OpenAI open weight models on AWS Bedrock to eight new regions. This expansion brings these powerful AI models closer to customers in various parts of the world, enabling lower latency and improved performance for a wide range of AI-powered applications. With this expansion, the OpenAI open weight models are now available in the following AWS Regions: US East (N. Virginia), Asia Pacific (Tokyo), Europe (Stockholm), Asia Pacific (Mumbai), Europe (Ireland), South America (São Paulo), Europe (London), and Europe (Milan), in addition to the previously supported region of US West (Oregon). This broader availability allows more customers to leverage these state-of-the-art AI models while keeping their data within their preferred geographic locations, helping to address data residency requirements and reduce network latency. To learn more about OpenAI open weight models on AWS Bedrock and how to get started, visit the Amazon Bedrock console or check out our documentation. For more information about the initial release of these models on AWS Bedrock, refer to our previous blog post.   

Publicado el Deja un comentario

Qwen3 models are now available fully managed in Amazon Bedrock

Amazon Bedrock continues to expand model choice by adding four Qwen3 open weight foundation models, now available as fully managed, serverless offerings. The lineup includes: Qwen3-Coder-480B-A35B-Instruct, Qwen3-Coder-30B-A3B-Instruct, Qwen3-235B-A22B-Instruct-2507, and Qwen3-32B for efficient dense computation. These models feature both dense and Mixture-of-Experts (MoE) architectures, providing flexible options for various development needs.

These open weight models enable you to build powerful AI applications with advanced agentic capabilities, without managing any infrastructure. The two Qwen3-Coder models excel at agentic coding and complex software engineering tasks, offering state-of-the-art performance for function calling and tool use. The 235B model delivers efficient general reasoning and instruction following across diverse tasks, while the 32B dense model provides a more traditional architecture suitable for a wide range of computational tasks.

Qwen3 models (32B, Coder-30B) are available today in the US East (N. Virginia), US West (Oregon), Asia Pacific (Mumbai, Tokyo), Europe (Ireland, London, Milan, Stockholm), and South America (São Paulo) AWS Regions. Qwen 235B is available today in theUS West (Oregon), Asia Pacific (Mumbai, Tokyo), and Europe (London, Milan, Stockholm) AWS Regions. Qwen Coder-480B is available today in the US West (Oregon), Asia Pacific (Mumbai, Tokyo), and Europe (London, Stockholm) AWS Regions. Check the full Region list for future updates. To learn more, read the blog, product page, Amazon Bedrock pricing, and documentation. To get started with Qwen in Amazon Bedrock, visit the Amazon Bedrock console.

 

​Amazon Bedrock continues to expand model choice by adding four Qwen3 open weight foundation models, now available as fully managed, serverless offerings. The lineup includes: Qwen3-Coder-480B-A35B-Instruct, Qwen3-Coder-30B-A3B-Instruct, Qwen3-235B-A22B-Instruct-2507, and Qwen3-32B for efficient dense computation. These models feature both dense and Mixture-of-Experts (MoE) architectures, providing flexible options for various development needs. These open weight models enable you to build powerful AI applications with advanced agentic capabilities, without managing any infrastructure. The two Qwen3-Coder models excel at agentic coding and complex software engineering tasks, offering state-of-the-art performance for function calling and tool use. The 235B model delivers efficient general reasoning and instruction following across diverse tasks, while the 32B dense model provides a more traditional architecture suitable for a wide range of computational tasks. Qwen3 models (32B, Coder-30B) are available today in the US East (N. Virginia), US West (Oregon), Asia Pacific (Mumbai, Tokyo), Europe (Ireland, London, Milan, Stockholm), and South America (São Paulo) AWS Regions. Qwen 235B is available today in theUS West (Oregon), Asia Pacific (Mumbai, Tokyo), and Europe (London, Milan, Stockholm) AWS Regions. Qwen Coder-480B is available today in the US West (Oregon), Asia Pacific (Mumbai, Tokyo), and Europe (London, Stockholm) AWS Regions. Check the full Region list for future updates. To learn more, read the blog, product page, Amazon Bedrock pricing, and documentation. To get started with Qwen in Amazon Bedrock, visit the Amazon Bedrock console.  

Publicado el Deja un comentario

DeepSeek-V3.1 model now available fully managed in Amazon Bedrock

DeepSeek-V3.1 is now available as a fully managed foundation model in Amazon Bedrock. This advanced open weight model allows you to switch between thinking mode for detailed step-by-step analysis and non-thinking mode for quicker responses. With comprehensive multilingual support, it delivers enhanced accuracy and reduced hallucinations compared to previous DeepSeek models, while maintaining visibility into its decision-making process.

You can use DeepSeek-V3.1’s enterprise-grade capabilities across critical business functions, from state-of-the-art software development to complex mathematical reasoning and data analysis. The model excels at sophisticated problem-solving tasks, demonstrating strong performance in coding benchmarks and technical challenges. Its enhanced tool-calling capabilities and seamless workflow integration make it ideal for building AI agents and automating enterprise processes, while its transparent reasoning approach helps teams understand and trust its outputs.

DeepSeek-V3.1 is now available in the US West (Oregon), Asia Pacific (Tokyo), Asia Pacific (Mumbai), Europe (London), and Europe (Stockholm) AWS Regions. To learn more, read the blog, product page, Amazon Bedrock pricing, and documentation. To get started with DeepSeek in Amazon Bedrock, visit the Amazon Bedrock console.

 

​DeepSeek-V3.1 is now available as a fully managed foundation model in Amazon Bedrock. This advanced open weight model allows you to switch between thinking mode for detailed step-by-step analysis and non-thinking mode for quicker responses. With comprehensive multilingual support, it delivers enhanced accuracy and reduced hallucinations compared to previous DeepSeek models, while maintaining visibility into its decision-making process. You can use DeepSeek-V3.1’s enterprise-grade capabilities across critical business functions, from state-of-the-art software development to complex mathematical reasoning and data analysis. The model excels at sophisticated problem-solving tasks, demonstrating strong performance in coding benchmarks and technical challenges. Its enhanced tool-calling capabilities and seamless workflow integration make it ideal for building AI agents and automating enterprise processes, while its transparent reasoning approach helps teams understand and trust its outputs. DeepSeek-V3.1 is now available in the US West (Oregon), Asia Pacific (Tokyo), Asia Pacific (Mumbai), Europe (London), and Europe (Stockholm) AWS Regions. To learn more, read the blog, product page, Amazon Bedrock pricing, and documentation. To get started with DeepSeek in Amazon Bedrock, visit the Amazon Bedrock console.  

Publicado el Deja un comentario

Microsoft 365 Copilot: Habilitación de equipos de agentes humanos

septiembre 18, 2025

Microsoft 365 Copilot: Habilitación de equipos de agentes humanos

Una mujer y un hombre sostienen unas laptops. al lado, un texto dice "Microsoft 365 Copilot"

Por: Nicole Herskowitz, vicepresidenta corporativa, Microsoft 365 y Copilot.

El trabajo es, de manera fundamental, un deporte de equipo, pero hasta ahora la IA ha sido en gran medida un asistente personal. Hoy, presentamos nuevos agentes centrados en la colaboración para los usuarios de Microsoft 365 Copilot, lo que brinda a cada equipo, proyecto, reunión y comunidad un compañero de equipo de IA, que agrega IA sensible al contexto para respaldar las necesidades únicas de cada escenario de colaboración.

Estos nuevos agentes colaborativos están diseñados para mejorar el trabajo en Microsoft Teams, SharePoint y Viva Engage, para ayudar a los grupos a coordinarse, comunicarse y ejecutar con mayor claridad y eficacia. Al aprovechar la inteligencia de trabajo de Microsoft Graph, estos agentes ofrecen soporte contextual al tiempo que mantienen los controles de seguridad, identidad, cumplimiento y administración de nivel empresarial, lo que ayuda a que las interacciones sigan productivas y protegidas. Lean a continuación para explorar cómo estos agentes se convierten en participantes activos en cada etapa del trabajo en equipo.

Habilitación de equipos humano-agente

Los nuevos agentes diseñados con propósito, proporcionan IA siempre activa integrada donde ocurre la colaboración. Cada agente se basa en el contexto del grupo y está equipado con habilidades adaptadas a las formas únicas en que los grupos colaboran en canales, reuniones y comunidades en Teams, y bibliotecas y sitios en SharePoint. Digamos, por ejemplo, que un equipo trabaja en el lanzamiento de un producto para “Project Pluto».

El equipo consolida todas sus conversaciones y planes en un canal dedicado al «Project Pluto» en Teams, que ahora está equipado con un «Agente del Project Pluto». Los usuarios del canal pueden dirigir a este agente para resumir hilos, destilar decisiones, redactar planes y publicaciones, programar puntos de control y coordinarse con el agente de Project Manager para crear tareas para mantener el trabajo en movimiento.

Cuando llega el momento de una reunión de planeación en Teams para discutir el «Project Pluto», el agente facilitador de las reuniones interviene para preparar las agendas. Durante la reunión, toma notas de manera proactiva, mantiene la discusión en el buen camino, captura las decisiones y las convierte en acciones propias con seguimientos, rastreadas por completo a través de la integración con el agente del gerente de proyectos, e incluso completa algunas tareas por su cuenta. Los participantes de la reunión pueden guiar de manera colectiva al agente para que haga cosas como reorganizar la agenda o establecer un temporizador de reunión. 

Y para amplificar el lanzamiento del producto en la «Comunidad de ventas» en Viva Engage, el «Agente de la comunidad de ventas» maneja los anuncios, responde preguntas comunes con fuentes citadas y ayuda a los administradores de la comunidad a mantener las discusiones activas y precisas en tiempo real.

Detrás de escena, el Agente de conocimiento en SharePoint mantiene en forma el espacio de trabajo del «Project Pluto». Organiza y enriquece archivos, aplica las etiquetas correctas, realiza un seguimiento de las actualizaciones y une contenido relacionado del canal de Teams, las reuniones y la comunidad de ventas. Por lo tanto, cuando alguien le hace una pregunta a Microsoft 365 Copilot, ya sea «¿Cuál es nuestro posicionamiento aprobado?» o «¿Qué especificación es definitiva?»—extrae la fuente autorizada con citas.

Juntos, estos agentes mantienen cada etapa del lanzamiento del «Project Pluto», desde la planificación hasta la ejecución y la comunicación, con un funcionamiento sin problemas, con la IA que trabaja junto al equipo. Teams también admite un ecosistema abierto de agentes creados por socios y, con Model Context Protocol (MCP), esos agentes pueden colaborar sin problemas con agentes nativos de Teams, para compartir contexto e invocar las herramientas de los demás dentro del mismo flujo de trabajo. Estamos entusiasmados de trabajar con socios para desarrollar soluciones para que la colaboración multiplataforma con Teams sea sencilla.

Reimaginar el trabajo en equipo con IA: Primeros pasos

Microsoft 365 Copilot va más allá de la productividad personal para permitir que los equipos trabajen juntos con IA, para crear estrategias, reducir la falta de comunicación y acelerar el progreso. Estos nuevos agentes de colaboración ya están disponibles para todos los usuarios de Microsoft 365 Copilot en versión preliminar pública, y Facilitador para reuniones de Teams ya está disponible con carácter general. Dado que estas experiencias se basan en los mismos estándares de seguridad, cumplimiento y privacidad que sustentan Microsoft 365, las organizaciones pueden adoptarlas con confianza.

Para comenzar, pidan ayuda a cualquiera de estos agentes dondequiera que trabajen. Aquí hay algunas formas de colaborar con IA:

  • Utilicen Facilitador para su próxima reunión de equipo para generar una agenda, capturar decisiones, asignar seguimientos en automático y probar la presión de los puntos principales planteados durante la reunión.
  • Pidan al agente de uno de sus canales de Teams que redacte un informe de estado de la semana pasada, basado en las conversaciones del canal y en las reuniones en las que participaron.
  • Habiliten un agente en su comunidad de Viva Engage más ocupada y comparen las respuestas proporcionadas por el agente de su comunidad con su respuesta.
  • Pidan al Knowledge Agent que etiquete y organice todos los archivos relevantes para un proyecto en el que trabajen, luego creen un resumen del proyecto que pueda compartir con las partes interesadas.

Desbloquear la colaboración de IA con Microsoft

Obtengan más información sobre Microsoft 365 Copilot y exploren los recursos siguientes para obtener más información sobre las funcionalidades de IA colaborativa descritas arriba:

The post Microsoft 365 Copilot: Habilitación de equipos de agentes humanos appeared first on Source LATAM.

 

​The post Microsoft 365 Copilot: Habilitación de equipos de agentes humanos appeared first on Source LATAM.  

Publicado el Deja un comentario

Second-generation AWS Outposts racks now supported in the AWS Canada (Central) and US West (N. California) Regions

Second-generation AWS Outposts racks are now supported in the AWS Canada (Central) and US West (N. California) Regions. Outposts racks extend AWS infrastructure, AWS services, APIs, and tools to virtually any on-premises data center or colocation space for a truly consistent hybrid experience.

Organizations from startups to enterprises and the public sector in and outside of Canada and the US can now order their Outposts racks connected to these two new supported Regions, optimizing for their latency and data residency needs. Outposts allows customers to run workloads that need low-latency access to on-premises systems locally while connecting back to their home Region for application management. Customers can also use Outposts and AWS services to manage and process data that needs to remain on-premises to meet data residency requirements. This regional expansion provides additional flexibility in the AWS Regions that customers’ Outposts can connect to.

To learn more about second-generation Outposts racks, read this blog post and user guide. For the most updated list of countries and territories and the AWS Regions where second-generation Outposts racks are supported, check out the Outposts racks FAQs page.

 

​Second-generation AWS Outposts racks are now supported in the AWS Canada (Central) and US West (N. California) Regions. Outposts racks extend AWS infrastructure, AWS services, APIs, and tools to virtually any on-premises data center or colocation space for a truly consistent hybrid experience.
Organizations from startups to enterprises and the public sector in and outside of Canada and the US can now order their Outposts racks connected to these two new supported Regions, optimizing for their latency and data residency needs. Outposts allows customers to run workloads that need low-latency access to on-premises systems locally while connecting back to their home Region for application management. Customers can also use Outposts and AWS services to manage and process data that needs to remain on-premises to meet data residency requirements. This regional expansion provides additional flexibility in the AWS Regions that customers’ Outposts can connect to.
To learn more about second-generation Outposts racks, read this blog post and user guide. For the most updated list of countries and territories and the AWS Regions where second-generation Outposts racks are supported, check out the Outposts racks FAQs page.  

Publicado el Deja un comentario

Amazon Lex provides enhanced confirmation and currency built-in slots to 10 additional languages

Amazon Lex now provides support for confirmation and currency slot types in 10 additional languages: Portuguese, Catalan, French, Italian, German, Spanish, Mandarin, Cantonese, Japanese, and Korean. Built-in slots help you build more natural and efficient conversations by understanding synonyms of what you user says and resolving those inputs to a standard format. The confirmation slot helps understand various expressions of user acknowledgement and converts them into ‘Yes’, ‘No’, “Don’t know’‘, or ‘Maybe’. The currency slot helps identify currency and represents the input in a structured way. For example, when a user says “nope” or “absolutely not”, the confirmation slot resolves to ‘No’ or when the user says “1 dollar’, the currency slot resolves it to ”USD 1.00“. These built-in slots help you build more natural and efficient conversational experiences.

This feature is available in all commercial AWS Regions where Amazon Lex operates. To learn more about these features, visit Amazon Lex documentation or to learn how Amazon Connect and Amazon Lex deliver cloud-based conversational AI experiences for contact centers, please visit the Amazon Connect website.

 

​Amazon Lex now provides support for confirmation and currency slot types in 10 additional languages: Portuguese, Catalan, French, Italian, German, Spanish, Mandarin, Cantonese, Japanese, and Korean. Built-in slots help you build more natural and efficient conversations by understanding synonyms of what you user says and resolving those inputs to a standard format. The confirmation slot helps understand various expressions of user acknowledgement and converts them into ‘Yes’, ‘No’, “Don’t know’‘, or ‘Maybe’. The currency slot helps identify currency and represents the input in a structured way. For example, when a user says “nope” or “absolutely not”, the confirmation slot resolves to ‘No’ or when the user says “1 dollar’, the currency slot resolves it to ”USD 1.00“. These built-in slots help you build more natural and efficient conversational experiences. This feature is available in all commercial AWS Regions where Amazon Lex operates. To learn more about these features, visit Amazon Lex documentation or to learn how Amazon Connect and Amazon Lex deliver cloud-based conversational AI experiences for contact centers, please visit the Amazon Connect website.  

Publicado el Deja un comentario

Amazon OpenSearch Serverless now supports Disk-Optimized Vectors

We are excited to announce the launch of disk-optimized vector support for Amazon OpenSearch Serverless, offering customers a cost-effective solution for vector search operations without compromising on accuracy and recall rates. This new feature enables organizations to implement high-quality vector search capabilities while significantly reducing operational costs.

With the introduction of Disk Optimized Vectors, customers can now choose between memory-optimized and disk-optimized vector storage options. The disk-optimized option delivers the same high accuracy and recall rates as memory-optimized vectors at lower cost. While this option may introduce slightly higher latency, it’s ideal for use cases where sub-millisecond response times aren’t critical such as semantic search applications, recommendation systems, and other AI-powered search scenarios.

Amazon OpenSearch Serverless, our fully managed deployment option, eliminates the complexities of infrastructure management for search and analytics workloads. The service automatically scales compute capacity, measured in OpenSearch Compute Units (OCUs), based on your workload demands.

Please refer to the AWS Regional Services List for more information about Amazon OpenSearch Service availability. To learn more about OpenSearch Serverless, see the documentation.

 

​We are excited to announce the launch of disk-optimized vector support for Amazon OpenSearch Serverless, offering customers a cost-effective solution for vector search operations without compromising on accuracy and recall rates. This new feature enables organizations to implement high-quality vector search capabilities while significantly reducing operational costs. With the introduction of Disk Optimized Vectors, customers can now choose between memory-optimized and disk-optimized vector storage options. The disk-optimized option delivers the same high accuracy and recall rates as memory-optimized vectors at lower cost. While this option may introduce slightly higher latency, it’s ideal for use cases where sub-millisecond response times aren’t critical such as semantic search applications, recommendation systems, and other AI-powered search scenarios. Amazon OpenSearch Serverless, our fully managed deployment option, eliminates the complexities of infrastructure management for search and analytics workloads. The service automatically scales compute capacity, measured in OpenSearch Compute Units (OCUs), based on your workload demands.
Please refer to the AWS Regional Services List for more information about Amazon OpenSearch Service availability. To learn more about OpenSearch Serverless, see the documentation.  

Publicado el Deja un comentario

AWS Step Functions now supports IPv6 with dual-stack endpoints

AWS Step Functions adds supports for IPv6. You can now send IPV6 traffic to AWS Step Functions via new dual-stack IPv4 and IPv6 endpoints. AWS Step Functions is a visual workflow service that enables customers to build distributed applications, automate IT and business processes, and build data and machine learning pipelines using AWS services. This enhancement addresses the growing need for IP addresses as the internet continues to expand, providing a larger address space than the traditional IPv4 format.

With IPv6 support, organizations modernizing their applications can now build serverless workflows without being constrained by limited IPv4 address space. The new dual-stack endpoints support both IPv4 and IPv6 protocols while maintaining backwards compatibility with existing IPv4 endpoints. Step Functions also supports IPv6 connectivity through PrivateLink interface Virtual Private Cloud (VPC) endpoints, enabling you to access the service privately without traversing the public internet. This enables organizations operating in IPv6 environments to natively integrate with Step Functions without requiring complex translation mechanisms between IPv6 and IPv4.

IPv6 support for AWS Step Functions is now generally available in US East (Ohio), US East (N. Virginia), US West (Oregon), US West (N. California) as well as AWS GovCloud (US-East), and AWS GovCloud (US-West) Regions, where AWS Step Functions is available.

To learn more about IPv6 support on AWS, visit the documentation page.

 

​AWS Step Functions adds supports for IPv6. You can now send IPV6 traffic to AWS Step Functions via new dual-stack IPv4 and IPv6 endpoints. AWS Step Functions is a visual workflow service that enables customers to build distributed applications, automate IT and business processes, and build data and machine learning pipelines using AWS services. This enhancement addresses the growing need for IP addresses as the internet continues to expand, providing a larger address space than the traditional IPv4 format. With IPv6 support, organizations modernizing their applications can now build serverless workflows without being constrained by limited IPv4 address space. The new dual-stack endpoints support both IPv4 and IPv6 protocols while maintaining backwards compatibility with existing IPv4 endpoints. Step Functions also supports IPv6 connectivity through PrivateLink interface Virtual Private Cloud (VPC) endpoints, enabling you to access the service privately without traversing the public internet. This enables organizations operating in IPv6 environments to natively integrate with Step Functions without requiring complex translation mechanisms between IPv6 and IPv4. IPv6 support for AWS Step Functions is now generally available in US East (Ohio), US East (N. Virginia), US West (Oregon), US West (N. California) as well as AWS GovCloud (US-East), and AWS GovCloud (US-West) Regions, where AWS Step Functions is available. To learn more about IPv6 support on AWS, visit the documentation page.  

Publicado el Deja un comentario

Amazon SageMaker HyperPod now supports autoscaling using Karpenter

Amazon SageMaker HyperPod now supports managed node autoscaling using Karpenter, enabling customers to automatically scale their clusters to meet dynamic inference and training demands. Real-time inference workloads require automatic scaling to address unpredictable traffic patterns and maintain service level agreements, while optimizing costs. However, organizations often struggle with the operational overhead of installing, configuring, and maintaining complex autoscaling solutions. HyperPod-managed node autoscaling eliminates the undifferentiated heavy lifting of Karpenter setup and maintenance, while providing integrated resilience and fault tolerance capabilities.

Autoscaling on HyperPod with Karpenter enables customers to achieve just-in-time provisioning that rapidly adapts GPU compute for inference traffic spikes. Customers can scale to zero nodes during low-demand periods without maintaining dedicated controller infrastructure and benefit from workload-aware node selection that optimizes instance types and costs. For inference workloads, this provides automatic capacity scaling to handle production traffic bursts, cost reduction through intelligent node consolidation during idle periods, and seamless integration with event-driven pod autoscalers like KEDA. Training workloads also benefit from automatic resource optimization during model development cycles. You can enable autoscaling on HyperPod using the UpdateCluster API with AutoScaling mode set to «Enable» and AutoScalerType set to «Karpenter».

This feature is available in all AWS Regions where Amazon SageMaker HyperPod EKS clusters are supported. To learn more about autoscaling on SageMaker HyperPod with Karpenter, see the user guide and blog.

 

​Amazon SageMaker HyperPod now supports managed node autoscaling using Karpenter, enabling customers to automatically scale their clusters to meet dynamic inference and training demands. Real-time inference workloads require automatic scaling to address unpredictable traffic patterns and maintain service level agreements, while optimizing costs. However, organizations often struggle with the operational overhead of installing, configuring, and maintaining complex autoscaling solutions. HyperPod-managed node autoscaling eliminates the undifferentiated heavy lifting of Karpenter setup and maintenance, while providing integrated resilience and fault tolerance capabilities. Autoscaling on HyperPod with Karpenter enables customers to achieve just-in-time provisioning that rapidly adapts GPU compute for inference traffic spikes. Customers can scale to zero nodes during low-demand periods without maintaining dedicated controller infrastructure and benefit from workload-aware node selection that optimizes instance types and costs. For inference workloads, this provides automatic capacity scaling to handle production traffic bursts, cost reduction through intelligent node consolidation during idle periods, and seamless integration with event-driven pod autoscalers like KEDA. Training workloads also benefit from automatic resource optimization during model development cycles. You can enable autoscaling on HyperPod using the UpdateCluster API with AutoScaling mode set to «Enable» and AutoScalerType set to «Karpenter». This feature is available in all AWS Regions where Amazon SageMaker HyperPod EKS clusters are supported. To learn more about autoscaling on SageMaker HyperPod with Karpenter, see the user guide and blog.