Publicado el Deja un comentario

Amazon ECS Managed Instances reduces GPU management fees by up to 60%

Amazon Elastic Container Service (Amazon ECS) Managed Instances now offers significantly reduced management fees for GPU and accelerated instance types. Beginning July 1, 2026, G-series ECS management fees are reduced by 35%, and P-series and AWS Trainium fees are reduced by 60%. These reductions apply automatically and no action is required from customers already using GPU instances with ECS Managed Instances.

With ECS Managed Instances, you get the application performance you want and the simplicity you need. Simply define your task requirements such as the number of vCPUs, memory size, and CPU architecture, and Amazon ECS automatically provisions, configures and operates most optimal EC2 instances within your AWS account using AWS-controlled access. You can also specify desired instance types, including GPU-accelerated, network-optimized, and burstable performance, to run your workloads on the instance families you prefer. ECS Managed Instances includes capabilities built specifically for accelerated workloads: GPU metrics (utilization, memory, and temperature) through Amazon CloudWatch Container Insights, and automatic health monitoring that detects GPU-specific hardware failures and replaces unhealthy instances to minimize workload disruption. With today’s pricing update, customers running GPU workloads on ECS Managed Instances can now benefit from fully managed infrastructure at lower management fees.

This pricing update is available in all AWS Regions where ECS Managed Instances is available. For the complete updated rate table, see ECS Managed Instances pricing. Amazon EKS is implementing identical management fee reductions for GPU instances on EKS Auto Mode. See the EKS What’s New Post for details. To learn more about ECS Managed Instances, visit the feature page, documentation, and AWS News launch blog.

 

​Amazon Elastic Container Service (Amazon ECS) Managed Instances now offers significantly reduced management fees for GPU and accelerated instance types. Beginning July 1, 2026, G-series ECS management fees are reduced by 35%, and P-series and AWS Trainium fees are reduced by 60%. These reductions apply automatically and no action is required from customers already using GPU instances with ECS Managed Instances. With ECS Managed Instances, you get the application performance you want and the simplicity you need. Simply define your task requirements such as the number of vCPUs, memory size, and CPU architecture, and Amazon ECS automatically provisions, configures and operates most optimal EC2 instances within your AWS account using AWS-controlled access. You can also specify desired instance types, including GPU-accelerated, network-optimized, and burstable performance, to run your workloads on the instance families you prefer. ECS Managed Instances includes capabilities built specifically for accelerated workloads: GPU metrics (utilization, memory, and temperature) through Amazon CloudWatch Container Insights, and automatic health monitoring that detects GPU-specific hardware failures and replaces unhealthy instances to minimize workload disruption. With today’s pricing update, customers running GPU workloads on ECS Managed Instances can now benefit from fully managed infrastructure at lower management fees. This pricing update is available in all AWS Regions where ECS Managed Instances is available. For the complete updated rate table, see ECS Managed Instances pricing. Amazon EKS is implementing identical management fee reductions for GPU instances on EKS Auto Mode. See the EKS What’s New Post for details. To learn more about ECS Managed Instances, visit the feature page, documentation, and AWS News launch blog.  

Publicado el Deja un comentario

Amazon EMR Serverless now supports larger worker sizes to run more compute and memory intensive workloads

Amazon EMR Serverless now offers larger worker configurations of 32 vCPUs with up to 244 GB of memory, allowing you to run more compute and memory-intensive workloads. Previously, the largest worker configuration available on EMR Serverless was 16 vCPUs with up to 120 GB of memory. Larger workers can help you improve the runtime performance as well as cost profiles for your workloads.

For shuffle-heavy workloads, larger workers reduce inefficient data transfers between executors. For jobs with data skew, larger workers reduce the chances of out-of-memory failures. For jobs that need to cache data, larger workers allow holding more data in memory, boosting job performance. To take advantage of these benefits, we recommend using larger workers for your compute and memory-intensive Spark and Hive workloads.

To learn more about different worker configurations, please visit EMR Serverless documentation. Larger workers are available in all AWS Regions where EMR Serverless is available.

 

​Amazon EMR Serverless now offers larger worker configurations of 32 vCPUs with up to 244 GB of memory, allowing you to run more compute and memory-intensive workloads. Previously, the largest worker configuration available on EMR Serverless was 16 vCPUs with up to 120 GB of memory. Larger workers can help you improve the runtime performance as well as cost profiles for your workloads.
For shuffle-heavy workloads, larger workers reduce inefficient data transfers between executors. For jobs with data skew, larger workers reduce the chances of out-of-memory failures. For jobs that need to cache data, larger workers allow holding more data in memory, boosting job performance. To take advantage of these benefits, we recommend using larger workers for your compute and memory-intensive Spark and Hive workloads.
To learn more about different worker configurations, please visit EMR Serverless documentation. Larger workers are available in all AWS Regions where EMR Serverless is available.  

Publicado el Deja un comentario

Amazon EC2 C8ine instances are now available in AWS Europe (Frankfurt) region

Starting today, Amazon Elastic Compute Cloud (Amazon EC2) C8ine instances are available in the AWS Europe (Frankfurt) region. C8ine instances are powered by custom sixth generation Intel Xeon Scalable processors, available only on AWS. These instances feature the latest sixth generation AWS Nitro cards, delivering up to 43% higher performance compared to previous generation C6in instances.

C8ine instances offer up to 2.5 times higher packet performance per vCPU versus prior generation network optimized instances, providing up to 2x higher network throughput for traffic going through Internet gateways compared to existing C6in network optimized instances. C8ine instances are designed for security and network virtual appliances, including virtual firewalls, load balancers, and Telco 5G UPF workloads.

Amazon EC2 C8ine instances are available in US East (N. Virginia), US West (Oregon), Asia Pacific (Tokyo), and Europe (Frankfurt) regions. C8ine instances are available via Savings Plans and On-Demand instances. For more information, visit the Amazon EC2 C8i instance pages.

 

​Starting today, Amazon Elastic Compute Cloud (Amazon EC2) C8ine instances are available in the AWS Europe (Frankfurt) region. C8ine instances are powered by custom sixth generation Intel Xeon Scalable processors, available only on AWS. These instances feature the latest sixth generation AWS Nitro cards, delivering up to 43% higher performance compared to previous generation C6in instances.
C8ine instances offer up to 2.5 times higher packet performance per vCPU versus prior generation network optimized instances, providing up to 2x higher network throughput for traffic going through Internet gateways compared to existing C6in network optimized instances. C8ine instances are designed for security and network virtual appliances, including virtual firewalls, load balancers, and Telco 5G UPF workloads.
Amazon EC2 C8ine instances are available in US East (N. Virginia), US West (Oregon), Asia Pacific (Tokyo), and Europe (Frankfurt) regions. C8ine instances are available via Savings Plans and On-Demand instances. For more information, visit the Amazon EC2 C8i instance pages.  

Publicado el Deja un comentario

Amazon EKS Auto Mode reduces GPU management fees by up to 60%

Amazon Elastic Kubernetes Service (Amazon EKS) Auto Mode now offers significantly reduced management fees for GPU and accelerated instance types. Beginning July 1, 2026, G-series Auto Mode management fees are reduced by 35%, and P-series and AWS Trainium fees are reduced by 60%. These reductions apply automatically to all EKS Auto Mode clusters and no action is required from customers already using GPU instances with Auto Mode.

EKS Auto Mode simplifies Kubernetes operations by automatically provisioning and managing infrastructure for machine learning inference, fine-tuning, rendering, and batch processing workloads. It includes capabilities built for accelerated workloads: automatic parallel image pulling and unpacking on GPU instances with local NVMe storage, so large container and model images start faster, and accelerator-aware node repair that detects GPU hardware failures and automatically replaces unhealthy nodes. With today’s price reduction, customers can run GPU workloads on Auto Mode at lower management fees, making its fully managed infrastructure more cost-effective.

This pricing update is available in all AWS Regions where EKS Auto Mode is available. Amazon ECS is implementing identical management fee reductions for GPU instances on ECS Managed Instances. See the ECS What’s New post for details.

To get started with GPU workloads on EKS Auto Mode, see the EKS for AI/ML documentation. For the complete updated rate table, see Amazon EKS pricing.

 

​Amazon Elastic Kubernetes Service (Amazon EKS) Auto Mode now offers significantly reduced management fees for GPU and accelerated instance types. Beginning July 1, 2026, G-series Auto Mode management fees are reduced by 35%, and P-series and AWS Trainium fees are reduced by 60%. These reductions apply automatically to all EKS Auto Mode clusters and no action is required from customers already using GPU instances with Auto Mode.
EKS Auto Mode simplifies Kubernetes operations by automatically provisioning and managing infrastructure for machine learning inference, fine-tuning, rendering, and batch processing workloads. It includes capabilities built for accelerated workloads: automatic parallel image pulling and unpacking on GPU instances with local NVMe storage, so large container and model images start faster, and accelerator-aware node repair that detects GPU hardware failures and automatically replaces unhealthy nodes. With today’s price reduction, customers can run GPU workloads on Auto Mode at lower management fees, making its fully managed infrastructure more cost-effective.
This pricing update is available in all AWS Regions where EKS Auto Mode is available. Amazon ECS is implementing identical management fee reductions for GPU instances on ECS Managed Instances. See the ECS What’s New post for details.
To get started with GPU workloads on EKS Auto Mode, see the EKS for AI/ML documentation. For the complete updated rate table, see Amazon EKS pricing.  

Publicado el Deja un comentario

AWS Security Hub adds impact analysis for exposure findings

Today, AWS Security Hub adds impact analysis to exposure findings, helping security teams understand the full scope of what an attacker could reach if an exposure is exploited. Impact analysis extends exposure findings by mapping the downstream resources that could be compromised beyond the initially exposed resource, giving teams deeper visibility into organizational risk.

Security Hub analyzes the effective permissions of IAM principals associated with exposed resources to identify privilege escalation paths to other resources in your account. The resulting scope of impact is displayed in the potential attack path graph, and a new Impact Assessment tab shows the prioritized chains of resources an attacker could traverse along with the specific permissions at each step. Security Hub factors the scope of impact into its severity scoring for exposure findings, and adjusts existing exposures as their scope of impact is identified or changes, so that exposures with greater downstream reach are prioritized appropriately.

To learn more, see Understanding exposure findings in the AWS Security Hub User Guide and the AWS Security Hub product page. For the full list of AWS Regions where Security Hub is available, see the AWS Regional Services List.

 

​Today, AWS Security Hub adds impact analysis to exposure findings, helping security teams understand the full scope of what an attacker could reach if an exposure is exploited. Impact analysis extends exposure findings by mapping the downstream resources that could be compromised beyond the initially exposed resource, giving teams deeper visibility into organizational risk. Security Hub analyzes the effective permissions of IAM principals associated with exposed resources to identify privilege escalation paths to other resources in your account. The resulting scope of impact is displayed in the potential attack path graph, and a new Impact Assessment tab shows the prioritized chains of resources an attacker could traverse along with the specific permissions at each step. Security Hub factors the scope of impact into its severity scoring for exposure findings, and adjusts existing exposures as their scope of impact is identified or changes, so that exposures with greater downstream reach are prioritized appropriately. To learn more, see Understanding exposure findings in the AWS Security Hub User Guide and the AWS Security Hub product page. For the full list of AWS Regions where Security Hub is available, see the AWS Regional Services List.  

Publicado el Deja un comentario

Amazon Cognito now supports self-service provisioned API rate limits

Amazon Cognito now allows you to increase or decrease your provisioned API rate limits on demand. Cognito has default rate limits for the maximum number of operations per second that you can perform in your user pools in each AWS Region, and you can purchase additional limits on adjustable API categories. With the new on-demand model, you can adjust your rate limits up or down more quickly to match your application’s traffic patterns.

Previously, to adjust your Cognito API rate limits, you would request an increase through Service Quotas, where requests are manually reviewed. This meant you had to plan rate limits in advance ahead of anticipated traffic spikes. Now, you have a new self-service experience to set your desired Cognito rate limit up to the account-level max limit using the Amazon Cognito console or the new limit provisioning API operations. Rate limit changes take effect immediately.

Self-service provisioned limits are available for adjustable API categories in all AWS Regions where Amazon Cognito is available. For pricing details of this add-on feature, see Amazon Cognito pricing page. To get started, see developer guide.

 

​Amazon Cognito now allows you to increase or decrease your provisioned API rate limits on demand. Cognito has default rate limits for the maximum number of operations per second that you can perform in your user pools in each AWS Region, and you can purchase additional limits on adjustable API categories. With the new on-demand model, you can adjust your rate limits up or down more quickly to match your application’s traffic patterns. Previously, to adjust your Cognito API rate limits, you would request an increase through Service Quotas, where requests are manually reviewed. This meant you had to plan rate limits in advance ahead of anticipated traffic spikes. Now, you have a new self-service experience to set your desired Cognito rate limit up to the account-level max limit using the Amazon Cognito console or the new limit provisioning API operations. Rate limit changes take effect immediately. Self-service provisioned limits are available for adjustable API categories in all AWS Regions where Amazon Cognito is available. For pricing details of this add-on feature, see Amazon Cognito pricing page. To get started, see developer guide.  

Publicado el Deja un comentario

Amazon SageMaker Studio now integrates with Hugging Face for one-click model deployment and customization

Amazon SageMaker Studio now supports direct integration from Hugging Face, letting you go from discovering a model to working with it inside a fully configured Studio environment in a single click. Select any supported model on Hugging Face and choose «Customize on SageMaker AI» or «Deploy on SageMaker AI» to land directly on the corresponding workflow page with the model pre-loaded and ready to use.

Previously, getting from model discovery to a working environment required navigating the AWS Console to find SageMaker AI, configuring an environment, setting up IAM permissions for serverless model customization, and in many cases requesting GPU quota increases through Service Quotas before running a first job. Now, new customers complete a standard AWS sign-up and receive a SageMaker Studio environment created in seconds with pre-configured permissions for serverless model customization jobs including fine-tuning with custom reward functions for reinforcement learning, model evaluation, and deployment to SageMaker or Bedrock endpoints. Verified customers receive default GPU access to G5, G6, and G4dn instances across endpoint deployments, training jobs, and notebooks without requesting quota increases, and quota limit and utilization information is visible for each instance type directly inside the Studio environment. Returning customers signing in from Hugging Face or SageMaker product pages select their environment and land directly inside SageMaker Studio with the model ready to use.

This feature is available in all AWS Commercial Regions where Amazon SageMaker Studio is supported. To get started, visit any supported model on Hugging Face and select «Customize on SageMaker AI» or «Deploy on SageMaker AI,» or click Get Started from the SageMaker Studio page. To learn more, see Service quotas for Studio in the Amazon SageMaker documentation.

 

​Amazon SageMaker Studio now supports direct integration from Hugging Face, letting you go from discovering a model to working with it inside a fully configured Studio environment in a single click. Select any supported model on Hugging Face and choose «Customize on SageMaker AI» or «Deploy on SageMaker AI» to land directly on the corresponding workflow page with the model pre-loaded and ready to use.
Previously, getting from model discovery to a working environment required navigating the AWS Console to find SageMaker AI, configuring an environment, setting up IAM permissions for serverless model customization, and in many cases requesting GPU quota increases through Service Quotas before running a first job. Now, new customers complete a standard AWS sign-up and receive a SageMaker Studio environment created in seconds with pre-configured permissions for serverless model customization jobs including fine-tuning with custom reward functions for reinforcement learning, model evaluation, and deployment to SageMaker or Bedrock endpoints. Verified customers receive default GPU access to G5, G6, and G4dn instances across endpoint deployments, training jobs, and notebooks without requesting quota increases, and quota limit and utilization information is visible for each instance type directly inside the Studio environment. Returning customers signing in from Hugging Face or SageMaker product pages select their environment and land directly inside SageMaker Studio with the model ready to use.
This feature is available in all AWS Commercial Regions where Amazon SageMaker Studio is supported. To get started, visit any supported model on Hugging Face and select «Customize on SageMaker AI» or «Deploy on SageMaker AI,» or click Get Started from the SageMaker Studio page. To learn more, see Service quotas for Studio in the Amazon SageMaker documentation.  

Publicado el Deja un comentario

Amazon SageMaker HyperPod now supports disaggregated prefill and decode

Amazon SageMaker HyperPod now supports Disaggregated Prefill and Decode (DPD), an inference optimization that separates the two phases of large language model (LLM) inference — prefill and decode — onto dedicated GPU pools and transfers the key-value (KV) cache between them over Elastic Fabric Adapter (EFA) using GPU-Direct RDMA. Customers running LLMs in production for chat assistants, agentic pipelines, retrieval-augmented generation, and long-document analysis need consistent per-token latency and predictable throughput under mixed traffic, but when prefill and decode share the same GPU, a single long-context request can stall token generation for every concurrent request and force customers to over-provision one phase to protect the other.

With DPD, customers run compute-bound prefill on one set of GPUs and memory-bandwidth-bound decode on another, so the two phases no longer contend for the same resources. This delivers more consistent per-token latency under sustained concurrency, higher goodput at strict latency SLOs, and the ability to scale prefill and decode capacity independently to match the input and output distribution of the workload. An intelligent router automatically directs long-context requests through the disaggregated path and sends shorter prompts directly to the decoder, so customers get the benefit on the traffic that needs it without paying transfer overhead on short prompts. Customers enable DPD by adding a `pdSpec` section to the same `InferenceEndpointConfig` custom resource they already use for inference endpoints on the HyperPod Inference Operator, and DPD is composable with the existing KV cache offloading and intelligent routing features on HyperPod.

DPD is available for SageMaker HyperPod clusters using the EKS orchestrator on EFA-capable instance types in all AWS Regions where Amazon SageMaker HyperPod is available. To learn more, see Disaggregated Prefill and Decode for HyperPod inference in the Amazon SageMaker AI Developer Guide.

 

​Amazon SageMaker HyperPod now supports Disaggregated Prefill and Decode (DPD), an inference optimization that separates the two phases of large language model (LLM) inference — prefill and decode — onto dedicated GPU pools and transfers the key-value (KV) cache between them over Elastic Fabric Adapter (EFA) using GPU-Direct RDMA. Customers running LLMs in production for chat assistants, agentic pipelines, retrieval-augmented generation, and long-document analysis need consistent per-token latency and predictable throughput under mixed traffic, but when prefill and decode share the same GPU, a single long-context request can stall token generation for every concurrent request and force customers to over-provision one phase to protect the other. With DPD, customers run compute-bound prefill on one set of GPUs and memory-bandwidth-bound decode on another, so the two phases no longer contend for the same resources. This delivers more consistent per-token latency under sustained concurrency, higher goodput at strict latency SLOs, and the ability to scale prefill and decode capacity independently to match the input and output distribution of the workload. An intelligent router automatically directs long-context requests through the disaggregated path and sends shorter prompts directly to the decoder, so customers get the benefit on the traffic that needs it without paying transfer overhead on short prompts. Customers enable DPD by adding a `pdSpec` section to the same `InferenceEndpointConfig` custom resource they already use for inference endpoints on the HyperPod Inference Operator, and DPD is composable with the existing KV cache offloading and intelligent routing features on HyperPod. DPD is available for SageMaker HyperPod clusters using the EKS orchestrator on EFA-capable instance types in all AWS Regions where Amazon SageMaker HyperPod is available. To learn more, see Disaggregated Prefill and Decode for HyperPod inference in the Amazon SageMaker AI Developer Guide.  

Publicado el Deja un comentario

Amazon EVS VCF 9.0 and 9.1 support

Today, we are announcing that Amazon Elastic VMware Service (EVS) now supports VMware Cloud Foundation (VCF) 9.0 and 9.1.

Amazon EVS lets you run the latest VCF software directly within your Amazon Virtual Private Cloud (VPC) on EC2 bare-metal instances. With this latest announcement, you now have complete control of the installation, operations, and management of the VMware virtualization solution running the VCF 9.0 and recently released VCF 9.1 versions. You can continue to use the same tools, processes, and skills on Amazon EVS that you use in your data center today, managing your VCF environment yourself or with an experienced AWS partner. With this, we’re also launching the Solutions for EVS GitHub repository with examples, templates, and infrastructure as code artifacts to help you get started.

This release is available in all regions where Amazon EVS is offered.

For more details, visit the launch blog, the Amazon EVS product detail page and user guide. 

 

​Today, we are announcing that Amazon Elastic VMware Service (EVS) now supports VMware Cloud Foundation (VCF) 9.0 and 9.1.
Amazon EVS lets you run the latest VCF software directly within your Amazon Virtual Private Cloud (VPC) on EC2 bare-metal instances. With this latest announcement, you now have complete control of the installation, operations, and management of the VMware virtualization solution running the VCF 9.0 and recently released VCF 9.1 versions. You can continue to use the same tools, processes, and skills on Amazon EVS that you use in your data center today, managing your VCF environment yourself or with an experienced AWS partner. With this, we’re also launching the Solutions for EVS GitHub repository with examples, templates, and infrastructure as code artifacts to help you get started.
This release is available in all regions where Amazon EVS is offered.
For more details, visit the launch blog, the Amazon EVS product detail page and user guide.   

Publicado el Deja un comentario

CloudWatch Application Signals now automatically captures errors, performance anomalies, and deployment events

Today, AWS announces Service Events for Amazon CloudWatch Application Signals, which automatically captures exception and latency event snapshots, function-level performance data, and deployment events from instrumented services without additional code changes. Customers can now quickly identify whether a deployment has introduced new exceptions by navigating to CloudWatch > Application Signals > [Service] > Errors in the CloudWatch console.

Service Events is available to any application with CloudWatch Application Signals enabled. Customers instrument their applications with the ADOT SDKs or the Amazon CloudWatch Observability EKS add-on. Once Application Signals is active, Service Events begins capturing exception and latency event snapshots and deployment events automatically. Optionally, customers can gain deeper performance visibility by turning on function-call metrics.

Service Events is available in all commercial AWS Regions. Supported languages are Java, Python, and JavaScript.

To get started, see Monitor service events in the Amazon CloudWatch User Guide. Service Events data is captured as logs. Function call metrics are captured as OpenTelemetry metrics. Standard CloudWatch pricing applies. For details, see CloudWatch pricing.

 

​Today, AWS announces Service Events for Amazon CloudWatch Application Signals, which automatically captures exception and latency event snapshots, function-level performance data, and deployment events from instrumented services without additional code changes. Customers can now quickly identify whether a deployment has introduced new exceptions by navigating to CloudWatch > Application Signals > [Service] > Errors in the CloudWatch console.
Service Events is available to any application with CloudWatch Application Signals enabled. Customers instrument their applications with the ADOT SDKs or the Amazon CloudWatch Observability EKS add-on. Once Application Signals is active, Service Events begins capturing exception and latency event snapshots and deployment events automatically. Optionally, customers can gain deeper performance visibility by turning on function-call metrics.
Service Events is available in all commercial AWS Regions. Supported languages are Java, Python, and JavaScript.
To get started, see Monitor service events in the Amazon CloudWatch User Guide. Service Events data is captured as logs. Function call metrics are captured as OpenTelemetry metrics. Standard CloudWatch pricing applies. For details, see CloudWatch pricing.