Intro
MongoDB performance tuning often starts with a vague complaint: queries are slow, the database is lagging, or CPU spikes at odd hours. This guide turns that complaint into a repeatable diagnostic workflow. You will learn how to inventory your environment, profile slow queries, design better indexes, adjust MongoDB server settings safely, and build a monitoring stack that catches regressions before users notice.
We focus on practical, version-scoped commands and realistic examples. Every step includes the command, expected output, and a rollback path. This is not a theoretical overview; it is a field manual for developers, DevOps engineers, and technical startup teams running MongoDB in production.
Version and Environment Inventory
Before touching any setting, know exactly what you are running. This inventory step prevents version mismatches and topologies that behave differently than expected.
Prerequisites
- Access to the MongoDB shell (
mongosh) or a driver connection. - Read permissions on the database and the
admindatabase for server status. - A terminal with network access to all mongod/mongos nodes.
Observe before changing Run the following read-only commands and capture the output:
// Connect to your deployment (replace placeholders)
mongosh "mongodb://<user>:<password>@<host>:<port>/admin"
// Server version and build info
db.version()
// Expected output (example): '7.0.14'
// Deployment topology: standalone, replica set, or sharded cluster
db.hello()
// Look for 'isWritablePrimary', 'setName', 'hosts' fields.
// Check current connections and operation counts
db.serverStatus().connections
// Example output: { "current" : 12, "available" : 51188, "totalCreated" : 3421 }
db.serverStatus().opcounters
// Example output: { "insert" : 10, "query" : 2500, "update" : 300, "delete" : 45, ... }
Expected signal vs. failure
- If
db.hello()showsisWritablePrimary: true, the node is acting as primary. In a replica set, confirm other members are inSECONDARYstate viadb.hello()on each. - If
connections.currentis near or equal toconnections.available, you are running out of connection capacity, which leads to refused connections. - If
opcounters.queryis dramatically higher than other operations over the same period, read-heavy workload tuning is warranted.
Blast radius and recovery This is fully read-only, so no recovery is needed. However, ensure you are connected to the correct node. If you accidentally connect to a production primary and run a heavy serverStatus() repeatedly, you add minor load. Use a read preference of secondaryPreferred for observation when possible.
Safe Configuration Path
After inventory, tune configuration one parameter at a time. This reduces risk and makes attribution of effects clear.
Prerequisites
- You have identified a specific performance symptom (e.g., high write latency, slow queries).
- You have a baseline measurement from your monitoring tool or
mongostat. - You know how to quickly revert the setting.
Example: Enabling the profiler for slow queries The database profiler collects information on operations that exceed a time threshold. In MongoDB 5.0+, use db.setProfilingLevel().
// List current profiling level
db.getProfilingStatus()
// Expected output: { "was" : 0, "slowms" : 100, "sampleRate" : 1.0 }
// Set profiling level 1 (log slow operations) with threshold 50 ms
db.setProfilingLevel(1, { slowms: 50 })
// Expected output: { "was" : 0, "slowms" : 100, "sampleRate" : 1.0, "ok" : 1 }
Verification After a few minutes or after reproducing the slow operation, query the system.profile collection on the database:
// Find slow queries in the last hour (sorted by duration descending)
db.system.profile.find({ ts: { $gt: new Date(Date.now() - 3600*1000) } }).sort({ millis: -1 }).limit(5).pretty()
Expected output includes the query predicate, plan summary, and millis field.
Recovery If profiling adds unacceptable overhead (usually minimal), disable it:
db.setProfilingLevel(0)
Verification and Diagnostics
Now you need to diagnose actual performance issues using read-only tools. This section covers essential diagnostic commands and how to interpret them.
1. Identify Slow Queries
If profiling is enabled, system.profile is the first stop. Alternatively, run currentOp() to see what is happening right now:
// Find active operations running longer than 100 ms
db.adminCommand({ currentOp: 1, active: true, secs_running: { $gt: 0.1 } })
Look for the op field (query, insert, update, etc.), the ns (namespace), and secs_running. A query op running long often signals a missing index or a bad plan.
2. Explain a Query
Use explain() to see how MongoDB executes a query. Example for a query on orders collection:
db.orders.find({ customer_id: 12345, status: "shipped" }).explain("executionStats")
Key fields:
executionStats.executionTimeMillis: total time in ms.executionStats.totalDocsExaminedvsexecutionStats.nReturned: iftotalDocsExaminedis much larger thannReturned, the query is scanning too many documents.queryPlanner.winningPlan.stage: stages likeCOLLSCANmean a full collection scan;IXSCANmeans an index was used.
Diagnostic conclusion If you see COLLSCAN and totalDocsExamined equals collection size but nReturned is small, you need an index.
3. Check Index Usage
db.orders.getIndexes()
This lists existing indexes. For performance diagnostics, also check serverStatus().metrics.queryExecutor for index scan vs collection scan counts.
4. Working Set and Cache
db.serverStatus().wiredTiger.cache
Important metrics:
bytes currently in the cache: should be close to the configured cache size under steady load.pages read into cachespikes indicate cache misses and disk I/O.
If cache misses are high, consider increasing wiredTiger.engineConfig.cacheSizeGB (see Settings Tuning below).
Index and Query Tuning
This is the most impactful performance tuning area. Bad indexing is the number one cause of slow MongoDB queries.
Creating the Right Index
Given the earlier orders query, create a compound index:
// Index on customer_id ascending, status descending
db.orders.createIndex({ customer_id: 1, status: -1 }, { name: "idx_customer_status" })
Verification Re-run explain("executionStats") on the query. Expected: winningPlan.stage becomes IXSCAN, totalDocsExamined drops dramatically, and executionTimeMillis decreases.
Index Selectivity and Order
Order matters in compound indexes. For queries that filter on both customer_id and status, the field that is most frequently used alone or has high selectivity should come first. Use db.orders.aggregate([{ $match: { customer_id: 12345, status: "shipped" } }, { $group: { _id: "$status", count: { $sum: 1 } } }]) to assess selectivity.
Avoiding Common Index Mistakes
- Over-indexing: Each index adds write overhead and consumes RAM. Only create indexes that support actual query patterns.
- Indexing low-cardinality fields: Indexing
statusalone (with few distinct values) often results in low selectivity and may not be used. - Ignoring sort stages: If a query sorts by
created_at, include that field in the index in the correct sort order to avoid in-memory sorting.
Monitoring index usage Use $indexStats to see how often each index is used:
db.orders.aggregate([{ $indexStats: {} }])
Look for indexes with accesses.ops of 0; they are candidates for removal (but confirm no hidden usage).
Server Settings Tuning
Beyond indexes, MongoDB server configuration can be adjusted. Always change one setting at a time and monitor.
WiredTiger Cache Size
For dedicated MongoDB servers, the default cache size is 50% of (RAM - 1 GB). This can be increased on memory-rich systems. You can set it in the config file:
storage:
wiredTiger:
engineConfig:
cacheSizeGB: 4
On a running server, use db.adminCommand({ setParameter: 1, wiredTigerEngineRuntimeConfig: "cache_size=4GB" }) but this does not persist across restart. Persist via config file or startup parameter.
Verification Check cache usage over time with db.serverStatus().wiredTiger.cache. Look for a decrease in pages read into cache and an increase in bytes currently in the cache up to the new limit.
Recovery If you see increased swapping or OOM issues, revert the cache size to its previous value.
oplog Size (Replica Sets)
The oplog must be large enough to hold operations between secondary sync cycles. If secondaries fall behind, they may enter RECOVERING state. Check current oplog size:
db.printReplicationInfo()
// Example output: configured oplog size: 990MB; log length start to end: 2.5hrs
If log length is less than your maintenance window, consider resizing. Changing oplog size requires a rolling restart with --oplogSize parameter or replSetResizeOplog (MongoDB 4.4+).
Connection Pool Limits
Drivers typically manage connection pools. Ensure the server's maxIncomingConnections is high enough:
db.adminCommand({ getParameter: 1, maxIncomingConnections: 1 })
Default is 65536, but you may need to lower it or raise it based on hardware. In the config file:
net:
maxIncomingConnections: 10000
Monitoring and Alerting
Proactive monitoring catches performance regressions early. Set up both database-level and host-level metrics.
Using mongostat and mongotop
mongostat provides a quick real-time view:
mongostat --host <host> --username <user> --password <password> --authenticationDatabase admin
Sample output columns: insert, query, update, delete, vsize, res, netIn, netOut, conn, time. Watch for high query or update rates and high conn (connections).
mongotop shows per-collection read/write activity:
mongotop --host <host> --username <user> --password <password> --authenticationDatabase admin
This helps identify hot collections.
Integration with Prometheus and Grafana
Use the mongodb_exporter (Percona or Bitnami) to scrape metrics. Key metrics to alert on:
mongodb_op_counters_totalfor operation rates.mongodb_mongod_connectionsfor connection count.mongodb_mongod_oplog_stats_sizefor oplog usage.mongodb_mongod_wiredtiger_cache_bytesfor cache usage.
Example Prometheus alert rule:
groups:
- name: mongodb_alerts
rules:
- alert: MongoDBHighConnections
expr: mongodb_mongod_connections > 5000
for: 10m
labels:
severity: warning
annotations:
summary: "High connections on {{ $labels.instance }}"
- alert: MongoDBCacheMissRate
expr: rate(mongodb_mongod_wiredtiger_cache_pages_read_into_cache[5m]) / rate(mongodb_mongod_wiredtiger_cache_pages_requested[5m]) > 0.05
for: 15m
labels:
severity: critical
annotations:
summary: "High cache miss rate on {{ $labels.instance }}"
Set up alerts for slow query count via profiling: you can export system.profile as logs or use a log shipper.
Common Pitfalls and How to Avoid Them
Even experienced engineers fall into these traps. Here is how to avoid or recover from each.
Pitfall 1: Adding Indexes Without Verifying
Why it happens: A developer sees a slow query and blindly adds an index suggested by a tool or blog post. The index may not match the query pattern, or it may cause write amplification.
How to avoid: Always run explain() before and after. Use the hint() method to force index usage in testing and compare performance. Check $indexStats after a few days to see if the index is actually used.
Recovery: If an index is unused or harmful, drop it with db.collection.dropIndex("index_name"). But first, confirm no application code expects it or uses it sporadically.
Pitfall 2: Ignoring Write Contention
Why it happens: Reads are optimized, but writes become slow due to too many indexes or document-level locks. Users complain about update latency.
How to avoid: Monitor db.serverStatus().wiredTiger.concurrentTransactions and look at write availability. Use fewer indexes on write-heavy collections. Consider sharding if write volume exceeds a single node's capacity.
Recovery: Remove unnecessary indexes and, if needed, redesign the schema to avoid large documents that cause page splits.
Pitfall 3: Setting Cache Size Too High
Why it happens: An operator sees high cache usage and increases cacheSizeGB to 80% of RAM, leaving little for the OS and causing swapping.
How to avoid: Follow the MongoDB documentation guideline: cacheSizeGB should be set to 50% of (RAM - 1 GB) or at most 80% of available memory for dedicated servers, but monitor host memory metrics. Use db.serverStatus().wiredTiger.cache and host free -m to ensure no swap.
Recovery: If you observe swap usage (vmstat 1 shows si or so non-zero), immediately reduce cacheSizeGB and restart mongod (or set at runtime and later persist).
Pitfall 4: Not Monitoring the Oplog
Why it happens: Small startup runs fine, then after a large data migration or high write period, secondaries fall behind and cannot catch up because the oplog has overwritten required entries.
How to avoid: Set up an alert for oplog window. Use db.printReplicationInfo() to check the window. The window should be at least the time needed for a full resync or maintenance.
Recovery: If a secondary is too far behind, perform an initial sync from a recent backup or another member. Increase oplog size proactively.
Pitfall 5: Neglecting Connection Pool Tuning in Application
Why it happens: Developers use default driver pool sizes, which may be too low for high concurrency or too high for limited server connections.
How to avoid: Benchmark with realistic concurrency. Set pool size to handle peak loads without overwhelming the server. In Node.js with the native driver:
const { MongoClient } = require('mongodb');
const client = new MongoClient(uri, { maxPoolSize: 100, minPoolSize: 10 });
Recovery: If you see frequent connection timeouts, increase pool size; if server connections are exhausted, reduce pool size or add connection limit on server.
Operations Checklist
Use this checklist before and after any performance change. Assign an owner to each item and review weekly or after each significant change.
| Item | Owner | Frequency | Verification |
|---|---|---|---|
| Baseline metrics recorded (latency, throughput, resource usage) | DBA Lead | Before change | Capture mongostat output and slow query count for 1 hour |
| Query patterns and index coverage analyzed | Backend Engineer | Before code deploy | explain() on top 10 queries shows IXSCAN, no COLLSCAN |
| New indexes created in staging and tested | Database Engineer | Before production rollout | Side-by-side performance comparison on clone |
| Application connection pool sized appropriately | Backend Engineer | Every release | Load test with expected peak concurrency |
| Server cache size and oplog size reviewed | DBA Lead | Monthly | db.printReplicationInfo() and db.serverStatus().wiredTiger.cache reviewed |
| Monitoring alerts enabled for key metrics | DevOps Engineer | Monthly | Test alert by simulating threshold breach |
| Rollback plan documented for each change | DBA Lead | Per change | Revert command or restore from backup verified |
Example owners: Priya Shah, Engineering Lead (DBA Lead); Alex Chen, Backend Engineer; Jordan Lee, DevOps Engineer.
Conclusion
MongoDB performance tuning is an iterative process: observe, hypothesize, change one variable, measure, and either keep or rollback. This guide has given you the tools to inventory your environment, find slow queries, add effective indexes, adjust server settings safely, and monitor for regressions.
Now, choose one low-risk action from this guide, such as enabling the profiler with a threshold of 50 ms, and run it in your development environment. Record the baseline, make the change, and compare the results after a day. The confidence you gain from this small, verified improvement will drive your broader performance initiatives.
Remember: the goal is not to apply every tuning tip, but to build a systematic approach that keeps your MongoDB deployment fast and stable.