Glossary
Capacity follows load automatically, so that neither users wait nor idle resources get paid for.
Scaling is driven by signals such as utilization, queue length or requests per instance. Choosing the signal matters: CPU is rarely what users feel.
Scaling down belongs to it as much as scaling up, and that is the part that saves money and gets forgotten more often.