Your AWS RDS bill is high. You log into Cost Explorer, see a number you were not expecting, and the breakdown tells you 'RDS,' but not why.
The problem is rarely one thing. Most organisations overpay on RDS because of a combination of causes that each look reasonable in isolation: an instance sized for peak load that never came, a read replica left running after the migration completed, a storage type purchased before gp3 existed, queries doing full table scans 400 times a minute.
This guide covers the seven causes that appear in most high RDS bills, with the specific diagnostic commands and SQL queries that identify each one.
Prefer to skip the audit and get the savings directly? Book a 30-min diagnostic ?
The Six Cost Categories in Your RDS Bill
Before diagnosing causes, it helps to know where your money is going.
Compute: the instance class running your database: db.r6g.2xlarge, db.t3.medium, and so on; typically the largest line item, often 50-65% of total RDS spend
Storage: the amount of data stored, in GB per month; pricing differs by storage type (gp2, gp3, io1, io2)
IOPS: on gp2, baseline IOPS are included (3 per GB, up to 16,000); on gp3 and io1/io2, IOPS are provisioned and priced separately
Backups and snapshots: automated backup storage beyond the free allocation (one times allocated storage), plus any manual snapshots you retain; these persist and accumulate cost indefinitely until deleted
Data transfer: data leaving RDS to other regions, to the internet, or to services in other availability zones
Multi-AZ and read replicas: a Multi-AZ standby runs a synchronous copy at full instance cost; each read replica is also billed at full instance cost
Log into AWS Cost Explorer, group by Usage Type, and identify which of these six categories is growing fastest. That tells you where to look first.
Related: a seventh cost that won't show up here: Extended Support fees on end-of-life engine versions
Cause #1: Your Database Is Overprovisioned
The most common cause of a high RDS bill is an instance sized for load that never materialised, or sized generously during a migration and never reviewed afterward. An instance running at 15-25% average CPU utilisation for 30 days is paying for 75-85% of its compute capacity to sit idle. 48/hour On-Demand, us-east-1), that is $345/month for two vCPUs doing useful work and eight vCPUs doing nothing.
24/hour) saves $173/month ($2,076/year), with no performance impact at that utilisation level. The diagnostic: pull 30-day average CPU utilisation and FreeableMemory from CloudWatch for every instance. If average CPU is below 30% and FreeableMemory consistently stays above 4 GB, the instance is a rightsizing candidate.
RDS supports instance class changes during the next maintenance window, or immediately with --apply-immediately and a brief connection interruption. Always test on a staging clone first and validate query latency under production-representative load before cutting over.
# Get 30-day average CPU utilisation for an RDS instance
aws cloudwatch get-metric-statistics \
--namespace AWS/RDS \
--metric-name CPUUtilization \
--dimensions Name=DBInstanceIdentifier,Value=prod-postgres \
--start-time $(date -u -d '30 days ago' '+%Y-%m-%dT%H:%M:%SZ') \
--end-time $(date -u '+%Y-%m-%dT%H:%M:%SZ') \
--period 2592000 \
--statistics Average \
--output table
# Get average FreeableMemory over 30 days (bytes � divide by 1073741824 for GB)
aws cloudwatch get-metric-statistics \
--namespace AWS/RDS \
--metric-name FreeableMemory \
--dimensions Name=DBInstanceIdentifier,Value=prod-postgres \
--start-time $(date -u -d '30 days ago' '+%Y-%m-%dT%H:%M:%SZ') \
--end-time $(date -u '+%Y-%m-%dT%H:%M:%SZ') \
--period 2592000 \
--statistics Average \
--output tableRelated: our AWS RDS performance tuning guide, the Database FinOps guide
Cause #2: Read Replicas Nobody Removed
Read replicas are created for good reasons: to offload read traffic, serve a reporting workload, or provide a low-lag source for analytics. The problem is that they rarely get removed when the original reason disappears. 24/hour adds $172/month regardless of how many queries it receives.
Replicas created before a migration to handle transition-period read traffic are forgotten. Replicas created for quarterly analytics jobs run at full cost 24/7 to serve four scheduled jobs per year. The diagnostic: list all replicas for every instance, then check DatabaseConnections in CloudWatch for each replica over the past 30 days.
A replica averaging fewer than 5 connections per day over 30 days, created more than 90 days ago, is a removal candidate. For workloads that only need occasional read offloading, consider creating the replica on demand and terminating it when the job completes. 44 per run rather than $172/month on standby.
# List all read replicas and their source instances
aws rds describe-db-instances \
--query 'DBInstances[?ReadReplicaSourceDBInstanceIdentifier!=null].{
ReplicaID:DBInstanceIdentifier,
SourceDB:ReadReplicaSourceDBInstanceIdentifier,
Class:DBInstanceClass,
Status:DBInstanceStatus,
Created:InstanceCreateTime
}' \
--output table
# Check connection activity on a replica over 30 days
# Low average connections = idle replica
aws cloudwatch get-metric-statistics \
--namespace AWS/RDS \
--metric-name DatabaseConnections \
--dimensions Name=DBInstanceIdentifier,Value=prod-postgres-replica \
--start-time $(date -u -d '30 days ago' '+%Y-%m-%dT%H:%M:%SZ') \
--end-time $(date -u '+%Y-%m-%dT%H:%M:%SZ') \
--period 86400 \
--statistics Average Maximum \
--output tableCause #3: Inefficient Queries Driving the Upgrade Spiral
The upgrade spiral is the most expensive pattern in cloud database costs. An endpoint slows down. The team adds a read replica.
The endpoint stays slow because the problem is query efficiency, not read volume. The team upgrades the instance class. The endpoint is still slow, and now the bill is 50% larger.
Another replica is added. Meanwhile, the root cause remains untouched: a query doing a full sequential scan on a 40-million-row table, 200 times per minute, generating 2 billion disk block reads per hour. High-I/O queries have two cost effects.
First, they directly generate IOPS charges: on gp3 with provisioned IOPS above the 3,000 baseline, every unnecessary disk read costs money. Second, they inflate CPU and memory utilisation, making an instance appear to be at capacity when it is doing unnecessary work. The diagnostic: enable pg_stat_statements and pull the top queries by total disk reads.
Any query with a cache miss percentage above 50% and more than 100 calls is generating avoidable I/O cost. Run EXPLAIN (ANALYZE, BUFFERS) on each to identify whether the fix is an index, a query rewrite, or a caching layer.
-- Enable pg_stat_statements
-- (requires parameter group change: shared_preload_libraries = 'pg_stat_statements')
CREATE EXTENSION IF NOT EXISTS pg_stat_statements;
-- Top queries by disk reads � the I/O cost drivers
SELECT
calls,
shared_blks_read AS disk_reads,
shared_blks_hit AS cache_hits,
round(100.0 * shared_blks_read /
nullif(shared_blks_read + shared_blks_hit, 0), 1) AS cache_miss_pct,
round(mean_exec_time::numeric, 1) AS avg_ms,
substring(query, 1, 120) AS query_short
FROM pg_stat_statements
WHERE calls > 100
ORDER BY shared_blks_read DESC
LIMIT 20;
-- Tables with active sequential scans (index candidates)
SELECT
schemaname,
relname AS table_name,
seq_scan,
idx_scan,
round(100.0 * seq_scan / nullif(seq_scan + idx_scan, 0), 1) AS seq_scan_pct,
n_live_tup AS live_rows
FROM pg_stat_user_tables
WHERE seq_scan > 100
ORDER BY seq_scan DESC
LIMIT 20;Related: how to debug slow SQL queries, the N+1 query problem
Still Piecing This Together Yourself?
A senior engineer looks at your actual setup, not a generic checklist, and tells you exactly what's wrong and how to fix it.
Book a Discovery CallCause #4: Storage Waste Accumulating Quietly
Storage charges on RDS accumulate in three places most teams do not actively monitor. First: manual snapshots. AWS creates automated backups on a retention schedule, which you pay for above the free allocation (one times allocated storage).
Manual snapshots, taken before deployments, before migrations, or at management request, persist indefinitely until explicitly deleted. 0575/GB/month in us-east-1). Twenty of them accumulated over two years cost $575/month, often invisible because the cost appears as 'RDS Snapshot Storage' rather than against any specific instance.
Second: backup retention set too long on non-production databases. Automated backup retention of 35 days on a 1 TB staging database costs roughly $34/month above the free allocation. Seven days is usually sufficient for staging: a $24/month saving per non-production instance.
Third: forgotten development and staging databases never decommissioned after a project ended. medium running 24/7 costs $46/month. Five forgotten dev instances accumulate to $276/month without serving any active workload.
To audit: go to RDS > Snapshots > Manual in the AWS console and review creation dates. Delete any manual snapshot older than 90 days with no active restoration use case.
Cause #5: Still on gp2 Storage
AWS launched gp3 storage for RDS in November 2021. If your instances were created before that date and were never reviewed, they are almost certainly on gp2, and paying a less favourable cost structure. 115/GB/month.
50/month difference on storage alone. More significantly, gp3 includes 3,000 IOPS and 125 MB/s throughput at no extra charge. On gp2, you get only 3 IOPS per GB: a 500 GB database gets 1,500 baseline IOPS.
065/IOPS-month above the baseline. 02/IOPS-month, 69% cheaper. The migration from gp2 to gp3 is live with no downtime.
Using aws rds modify-db-instance, the storage type changes while the database continues to serve traffic. No maintenance window, no reboot, and no backup-and-restore required.
# Find all RDS instances still on gp2 storage
aws rds describe-db-instances \
--query 'DBInstances[?StorageType==`gp2`].{
ID:DBInstanceIdentifier,
Class:DBInstanceClass,
StorageType:StorageType,
AllocatedGB:AllocatedStorage,
Iops:Iops
}' \
--output table
# Migrate from gp2 to gp3 � live migration, no downtime
# 3,000 IOPS and 125 MB/s throughput are included at no extra cost on gp3
aws rds modify-db-instance \
--db-instance-identifier prod-postgres \
--storage-type gp3 \
--iops 3000 \
--storage-throughput 125 \
--no-apply-immediatelyCause #6: On-Demand Pricing for Predictable Workloads
Every production RDS instance running 24/7 on On-Demand pricing is paying 30-50% more than it needs to. Reserved Instances commit you to a 1-year or 3-year term in exchange for a fixed hourly rate: the instance itself does not change, only the billing model. 48/hour On-Demand ($3,513/month for a single instance).
296/hour ($2,173/month). That is $1,340/month, or $16,080/year, saved on a single instance. 195/hour ($1,426/month), a 59% reduction.
The break-even for a 1-year reservation is roughly 7 months: if the instance will run for more than 7 months, the Reserved Instance costs less in total. The most common reason teams do not buy Reserved Instances is uncertainty about instance sizing: they do not want to lock in the wrong size. The right sequence: rightsize first (Cause #1), validate under production load for 30 days, then purchase the reservation.
Buying a 1-year commitment on an overprovisioned instance locks in the waste alongside the discount.
Related: managed DBA vs hiring in-house
Get a Straight Answer, Not a Sales Pitch
Tell us what you're running into. We'll tell you directly what's causing it and what it takes to fix, before you sign anything.
Book a Discovery CallCause #7: Solving the Wrong Problem, and a 15-Minute Audit
The most expensive RDS mistake is applying an infrastructure solution to an efficiency problem. Missing indexes cause high CPU and slow queries. The team adds a read replica: the bill doubles and the problem remains.
The instance is upgraded: the bill increases 50% and the problem shifts slightly. Another replica is added and the bill doubles again. Throughout, a $0 index addition would have eliminated the root cause.
Before any infrastructure change, ask: is this a capacity problem or an efficiency problem? Capacity problems (peak write volume exceeding single-instance throughput, read volume genuinely saturating a properly indexed query set) warrant infrastructure changes. Efficiency problems (sequential scans, N+1 queries, unindexed foreign keys, reporting queries running on the production primary) require code and schema fixes.
Pull 30-day average CPU per instance: below 30% is a rightsizing candidate
Check FreeableMemory: above 30% of instance RAM consistently suggests memory is oversized
List all replicas and check DatabaseConnections for each: below 5 per day average over 30 days is idle
Pull pg_stat_statements top 20 by shared_blks_read: any query with cache_miss_pct above 50% needs investigation
Check storage type: gp2 is a migration candidate
Check backup retention on non-production instances: above 14 days is review-worthy
Check pricing model: On-Demand on production instances that run continuously is a Reserved Instance candidate
A SaaS company paying $6,500/month ran this audit in 30 minutes. Findings: primary instance at 22% average CPU; one replica averaging 2 DatabaseConnections per day over 30 days; two queries doing sequential scans on a 30-million-row events table, running 150 times per minute; 47 retained manual snapshots from the previous 18 months; all three instances on On-Demand pricing. After rightsizing the primary, removing the idle replica, adding two covering indexes, deleting retained snapshots, and purchasing 1-year reservations, the monthly bill dropped from $6,500 to $4,200. Annual saving: more than $27,000.
-- 15-minute RDS cost audit � run on your PostgreSQL database
-- 1. Largest tables (identify where storage and I/O are concentrated)
SELECT
schemaname,
relname AS table_name,
pg_size_pretty(pg_total_relation_size(relid)) AS total_size,
pg_size_pretty(pg_relation_size(relid)) AS table_size,
pg_size_pretty(pg_indexes_size(relid)) AS index_size
FROM pg_catalog.pg_statio_user_tables
ORDER BY pg_total_relation_size(relid) DESC
LIMIT 10;
-- 2. Unused indexes (consuming storage and slowing writes for zero read benefit)
SELECT
schemaname,
relname AS table_name,
indexrelname AS index_name,
idx_scan AS times_used,
pg_size_pretty(pg_relation_size(indexrelid)) AS index_size
FROM pg_stat_user_indexes
WHERE idx_scan = 0
AND schemaname NOT IN ('pg_catalog', 'pg_toast')
ORDER BY pg_relation_size(indexrelid) DESC
LIMIT 20;
-- 3. Top I/O consumers (requires pg_stat_statements)
SELECT
calls,
round(total_exec_time::numeric / 1000, 2) AS total_cpu_sec,
shared_blks_read AS disk_reads,
round(100.0 * shared_blks_read /
nullif(shared_blks_read + shared_blks_hit, 0), 1) AS cache_miss_pct,
substring(query, 1, 100) AS query_short
FROM pg_stat_statements
WHERE calls > 50
ORDER BY shared_blks_read DESC
LIMIT 10;Most high RDS bills are not caused by growth. They are caused by drift. Instances sized at launch for capacity headroom that was never reclaimed.
Read replicas created for temporary workloads and never removed. Storage types not reviewed since the database was first provisioned. On-Demand pricing on instances that have been running for two years.
These costs accumulate month over month because no one is looking at the right metrics. The seven causes in this guide (overprovisioning, idle replicas, inefficient queries, storage waste, gp2 storage, On-Demand pricing, and solving the wrong problem) account for the majority of high RDS bills across organisations of all sizes. The 15-minute audit in Cause #7 covers all seven.
The starting point: pull 30-day CPU utilisation metrics per instance, list all read replicas and check their connection activity, and run the pg_stat_statements query to identify your top disk readers. Most of the savings are visible within the first 30 minutes of looking at the right data.
Frequently Asked Questions
Why is my AWS RDS bill so high?
The most common causes: overprovisioned compute running at 10-25% average CPU; read replicas left running after their original workload changed; storage on gp2 paying for IOPS that gp3 includes free; On-Demand pricing on production instances that qualify for 30-50% savings with Reserved Instances; and inefficient queries doing full table scans generating IOPS charges on every execution. Pull 30-day CPU utilisation metrics per instance and review your top queries by disk reads from pg_stat_statements. Most high RDS bills trace to one or more of these five causes.
How do I reduce AWS RDS costs?
Six changes cover the majority of RDS cost reductions: (1) rightsize instances running below 30% average CPU: a one-tier downgrade typically saves 30-50% on compute; (2) delete idle read replicas, or schedule them to exist only when the workload runs; (3) migrate gp2 storage to gp3: no downtime required, 20-25% storage cost reduction with better IOPS economics; (4) purchase 1-year Reserved Instances for production instances that run 24/7: saves 30-50% over On-Demand pricing; (5) add indexes to eliminate full table scan queries generating excess IOPS; (6) audit and delete retained manual snapshots and reduce backup retention on non-production databases from 35 days to 7.
What is usually the biggest RDS expense?
Compute (the instance class) is typically the largest single line item, accounting for 50-65% of total RDS spend. For most production workloads, storage and IOPS follow at 15-25%, and Multi-AZ or read replica compute adds another 15-25%. If compute is oversized, it inflates every downstream cost: Multi-AZ doubles the compute line, and each read replica adds another full instance at the same class price.
Should I use gp2 or gp3 for RDS storage?
gp3 for almost every use case. gp3 costs $0.092/GB/month versus gp2's $0.115/GB/month. gp3 includes 3,000 IOPS and 125 MB/s throughput at no extra cost. On gp2, you get only 3 IOPS per GB: a 500 GB database gets 1,500 baseline IOPS, and exceeding that costs $0.065/IOPS-month. On gp3, additional IOPS above 3,000 cost $0.02/IOPS-month, 69% cheaper. The migration from gp2 to gp3 is live with no downtime using aws rds modify-db-instance.
Are read replicas expensive on RDS?
Yes. Each read replica runs as a full RDS instance at the same class price as the source database. A db.r6g.xlarge replica at $0.24/hour adds $172/month whether it serves one query or a million. An idle replica costs the same as an active one. For workloads that only need occasional read offloading, such as quarterly reports or periodic analytics jobs, creating the replica on demand and terminating it after the job is far cheaper than running it continuously.
Can query optimization reduce my RDS bill?
Yes, and it often produces the largest savings relative to effort. A query doing a full sequential scan on a 30-million-row table generates hundreds of disk block reads per execution. At 100 executions per minute, that is millions of IOPS per day, each generating charges above the gp3 baseline or inflating CPU to the point where an upgrade seems necessary. Adding a covering index to eliminate the sequential scan simultaneously reduces IOPS charges, lowers CPU utilisation, and may enable a compute downsize. A single index can produce multiple cost reductions at once.
How much can rightsizing save on AWS RDS?
Rightsizing typically saves 20-50% on the compute component of an RDS bill. Since compute is 50-65% of most RDS bills, a 40% compute reduction translates to a 20-30% total bill reduction. A db.r6g.2xlarge ($0.48/hour On-Demand) downsized to db.r6g.large ($0.24/hour) saves $173/month on that single instance. Three production instances at the same ratio save $519/month ($6,228/year) before Reserved Instance savings are applied. Combined with Reserved Instance pricing after rightsizing, the total savings on compute can reach 60-70%.
Get a Straight Answer on Your Setup
Tell us what you're running into. We read every message personally and reply within 24 hours with times for a free call.


