Rightsizing Cloud Workloads on AMD, A Four-Part Series: Foundations
Oct 01, 2026
Cloud infrastructure rarely stays static. Over time, teams upgrade to newer instance generations to access faster CPUs, higher memory bandwidth, and better overall performance. These upgrades are often driven by availability, co-location agreements or reserved instances of end-of-life timelines or lift & shift cloud migrations. Most cloud migrations begin as a "lift and shift"; a like-for-like transfer of on-premises resources to the cloud. For example, taking your existing 24 core instance and simply moving it to the latest generation 48 vCPU instance in cloud. In practice, this is where many right-sizing opportunities are missed.
Rightsizing is commonly framed as a cost-cutting exercise, but in practice, it is the ongoing process of matching instance types and sizes to your workload performance and capacity requirements at the lowest possible cost. The goal is not simply to run on smaller instances, but to ensure that each workload is running on the right infrastructure for its behavior, performance requirements, and growth projections.
For example, on AWS, leveraging AMD CPU-powered instances (such as the M8a, C8a, and R8a) provides a unique opportunity to optimize. By understanding the specific topology of the AMD Core Complex Die (CCD), infrastructure engineers can do more than just lower a bill, they can improve application performance while spending less.
Each new instance generation of EPYC processors typically delivers meaningful architectural improvements, like, more powerful cores, better IPC (instructions per cycle), better L3 cache hierarchies, higher per-core performance, and/or improved efficiency. Instead of asking “Which new instance should we move to?”, the more effective question is "How much compute does this workload actually need on the newer architecture?". Despite these potential savings, many companies struggle to implement effective rightsizing. Why? Because rightsizing isn't just a one-time technical exercise, it requires a deep understanding of your workloads, clear business objectives, systematic analysis, and ongoing optimization as your applications & infrastructure evolves.
In this guide, we will move past basic utilization metrics. We’ll explore a holistic framework for rightsizing that considers your workload characterization, organization motivation, and the strategic advantage of aligning your compute footprint with AMD high-performance processor architecture. We’ll close with operational consideration for storage, containers, licensing, tagging, capacity planning, and usage commitments. Teams can turn recommendations into savings without performance regressions.
Throughout this series, we use AWS and AMD-powered EC2 instances as our primary examples because that is where many of these optimizations are easiest to illustrate. Where it is relevant, we note the equivalent capabilities on Microsoft Azure and Google Cloud, since the underlying rightsizing principles apply across all three.
Motivation and Ownership
Before analyzing telemetry or testing new architectures, it is important to answer a more fundamental question: Why are you rightsizing? How do you measure success?
Determining your primary motivation dictates how much risk you are willing to take and how much effort you should invest.
- Cost Optimization: For organizations facing immediate budget pressure to reduce cloud spending or meet FinOps targets. In these cases, the goal is often to achieve a measurable reduction. For example, if your primary goal is a fast 10–15% reduction in spend, you should focus on "low-hanging fruit." Downsizing idle or severely underutilized instances without deep application profiling. 3rd Gen AMD EPYC (formerly codenamed “Milan”) instances, 6a in AWS, are an optimally price value.
- Performance Maximization: If your application is hitting a performance ceiling, like latency issues, throughput constraints, or inconsistent behavior under load. Rightsizing focuses less on immediate savings and more on ensuring the workload is properly aligned with modern compute capabilities. By upgrading to the latest AMD instance families (like the M7a or M8a), you can leverage newer “Zen” architectures to address latency issues while keeping your budget flat.
- Price-Performance Optimization: This is the main goal for mature teams. It involves benchmarking your application to find the sweet spot where you get the maximum throughput for every dollar spent.
Regardless of the initial motivation, rightsizing almost always results in better price-performance. However, being explicit about the primary objective helps avoid misaligned expectations and unnecessary rework.
Ownership plays a significant role in how rightsizing is executed. In many enterprises, rightsizing is driven top-down by a centralized FinOps or cloud governance team. These teams analyze usage patterns, evaluate instance alternatives, including newer AMD EPYC™ CPU-based instance families, and provide recommendations to application owners. This model works well for broad cost optimization initiatives and generation upgrades. Other organizations operate in a bottom-up model, where application teams own both performance and infrastructure decisions. These teams are often better positioned to test alternative instance sizes, evaluate performance on AMD CPU-based instances, and fine-tune configurations based on real workload behavior.
A successful rightsizing culture requires a bridge between the FinOps team and Application Engineering. A hybrid approach is common in larger organizations, where central teams provide tooling, benchmarking guidance, and guardrails, while application teams that understand their unique workload, decide how and when to apply rightsizing changes. Teams could use tools like EPYC advisory services for instance recommendations and uProf to profile individual applications.
Timing of your Rightsizing
The timing of the rightsizing effort is also important. There isn’t only one “best” time, however, there are a few best moments, depending on your goal and how much risk and effort you can take on.
- Pre Migration Phase: The assessment and mobilize phases of a migration is usually the highest ROI timing because you can avoid carrying oversized infrastructure into cloud for months. This requires customers to commit time upfront and analyze infrastructure before the migration happens. The challenge is that sometimes on-premises monitoring data doesn’t predict cloud behavior, and heavily on-prem customers might lack expertise in cloud-native tools.
- During the Migration: Rightsizing can also be effective during migration itself, particularly when organizations migrate early pilot applications. These initial workloads should be low complexity with few dependencies, typically starting with non-production environments. Teams learn in the actual cloud environment rather than relying on synthetic performance benchmarks, validate their sizing assumptions against real performance data of their applications. This is good timing for wave-based migrations, where organizations are building cloud infra for the first time.
- After lift-and-shift: Many teams intentionally delay rightsizing to prioritize speed and stability. If your primary goal is moving applications quickly and safely, then rightsizing post migration to the cloud is efficient within the target cloud environment. While this approach can increase short-term cost, it is often the safest option for organizations with aggressive timelines or limited visibility in on prem environments.
- Ongoing operational process: If your workloads are already in cloud environments, rightsizing is most effective when treated as a continuous process rather than a one-time exercise. Your application behavior changes over time as traffic patterns evolve, code paths shift, and new features are introduced. You could establish a recurring cadence, quarterly or half yearly, allowing teams to reclaim unused capacity, and adapt to new infrastructure as workloads mature. This model aligns well with FinOps practices.
- Instance generation refreshes: Instance generation refreshes represent one of the lowest-effort, highest-return opportunities for rightsizing. When new infrastructure generations are announced or enter preview, they often deliver higher per-core performance, improved efficiency, and better overall price performance. Since each core is now more capable, you likely no longer need the same vCPU count to achieve your performance targets.
- Event-Driven Rightsizing: Rightsizing opportunities also arise in response to external or business-driven changes. Shifts in traffic patterns, new feature launches, regional expansions, or changes in service-level requirements can alter workload behavior.
Understanding Your Current State and Representative Benchmarks
Before you can select an instance, you must classify your application based on how it consumes resources and interacts with the underlying processor. The first step in deep discovery is workload classification. The same instance size can be overkill for a batch workload and insufficient for an always-on service. Without understanding runtime behavior, teams often apply the wrong optimization strategy. Different workload types stress different parts of the system. As a result, the choice of benchmark and the interpretation of its results are important tools for rightsizing.
Application Services: Application workloads typically include web services, APIs, and microservices that respond to user or system requests. These represent the broadest category, are often always-on, and performance typically depends on per-core efficiency, instruction throughput, and predictable scaling behavior rather than peak burst capacity. The goal is usually to find the highest and consistent performance at the lowest sustainable cost.
Because their performance is tightly coupled to the processor's raw speed and efficiency, SPEC CPU® (SPECrate®2017_int_base) benchmarking provides a normalized way to see how a newer generation like the AMD CPU-powered M8a (Turin) can handle significantly more work per vCPU compared to other instances. It provides a strong baseline for understanding relative CPU capability. However, SPEC CPU and SPECjbb workloads should be treated as application proxies for cloud instance evaluation and rightsizing. Results are intended for relative comparisons within this analysis only and are not directly comparable to published SPEC benchmark results.
Additional microbenchmarks are often used to represent specific application behaviors:
- SPECjbb® is commonly referenced for Java®-based services and JVM-heavy application stacks. Java applications are sensitive to memory latency and garbage collection pauses. This benchmark will give you an indicationof how the large L3 cache of an AMD CCD may help maintain high “critical-jOPS” and "max-jOPS[PB1.1]" (jOPS = Java operations per second)
- OpenSSL benchmarks help represent cryptographic workloads such as TLS termination and secure API traffic.
- NGINX benchmarks are frequently used to evaluate web-serving and request-handling throughput.
- CoreMark and 7-Zip are sometimes used to illustrate instruction efficiency and compression-heavy application paths.
On AWS, instances powered by AMD EPYC processors, like M7a and M8a, each vCPU maps to a full physical core, which provides predictable per-core performance and enables application workloads to run efficiently at higher utilization levels. These benchmarks help teams determine whether these instances can sustain the same throughput with fewer vCPUs[PB2.1]. In these cases, you could do either vertical or horizontal rightsizing or safely increasing autoscaling thresholds[PB3.1].
Databases and In-Memory Stores: Databases are often memory-bound rather than CPU-bound, meaning they benefit more from additional RAM than additional CPU cores beyond a certain point. They typically scale vertically rather than horizontally, though modern databases increasingly support read replicas and sharding for horizontal scaling. Database performance is about more than just raw clock speed, it also depends on memory bandwidth and L3 cache access. Performance characteristics vary widely depending on whether the system is transactional, analytical, or data is held in-memory. Software licensing is another variable to consider to maximize TCO by reducing core count. In many rightsizing exercises, teams use TPC-derived or workload-inspired test environments rather than full TPC-compliant benchmark implementations. Such results can be useful for comparative analysis but are not comparable to published TPC benchmark results.
- For Transactional (OLTP) workloads like MySQL or PostgreSQL, use TPC-C®-derived workloads[PB4.1]. For Microsoft SQL Server, you can use TPC-E derived workloads. These simulate complex transactional environments (orders, payments, inventory). Modern AMD instances, like the R8a, run with SMT OFF (Simultaneous Multi Threading). You will may see more consistent transaction latencies compared to hyper-threaded instances where two threads compete for a single core’s resources.
- For Analytical (OLAP) or data warehousing, use TPC[PB5.1]-H derived workloads. This focuses on complex queries and scan-heavy access patterns. This stresses long-running, complex queries and sequential throughput.
- For In-Memory stores like Redis, the redis-benchmark utility.
In addition to performance benchmarks, databases must include licensing analysis. On AWS, specifically latest 8th gen (8a) and 7th gen (7a) instances, EPYC CPU-based instances provide one physical core per vCPU, which can help reduce licensing costs for databases licensed per shared core. For example, an instance with eight vCPUs on M8a provides eight physical cores, reducing licensing costs by half compared to environments where SMT is enabled, where you might need 16 vCPU for same performance. In many enterprise deployments, database licensing costs exceed infrastructure costs making this a critical factor in rightsizing decisions.
Batch and Runtime-Based Workloads: These workloads are often not always-on and are measured by how fast they complete rather than how consistently they respond. For these workloads, rightsizing is closely tied to time-to-completion. Faster CPUs, higher memory bandwidth, and improved vector performance can directly reduce runtime. This allows teams to either finish jobs sooner or reduce the amount of infrastructure required. In this context, CPU performance is often the dominant factor, making SPEC CPU a useful representative or proxy benchmark. Higher SPEC CPU performance generally correlates with reduced job runtime, enabling teams to either downsize instances or reduce the number of instances required to complete the same work.
Conclusion
Strong rightsizing starts long before you pick an instance. In this first post we covered the foundations, knowing why you are rightsizing and how you will measure success, deciding who owns the effort, timing it for the highest return, and classifying your workloads so you choose benchmarks that actually represent how your applications behave. Get these right and every later decision becomes easier and lower risk.
In Part 2, we build on this foundation and go deeper into the AMD EPYC architecture itself, how the chiplet and CCD design shapes instance selection, and how to apply vertical and horizontal rightsizing strategies to match your compute footprint to real workload behavior.
Don’t Go It Alone
Rightsizing decisions touch performance, cost, licensing, and resilience all at once, and you do not have to navigate those trade-offs alone. If you want help classifying your workloads, choosing the right benchmarks, or building a rightsizing practice that sticks, reach out to AMD. We work directly with customers to help validate performance and map workloads to the right EPYC platforms.
Whether you are modernizing existing workloads, downsizing oversized environments, or simply looking for better performance with cost efficiency, AMD offers the tools and expertise to guide those decisions. For feedback on this post or to suggest future topics, reach out to AMD.
SPEC CPU®, SPECrate® and SPECjbb® are registered trademarks of Standard Performance Evaluation Corporation. Learn more at SPEC.org.