Backend engineeringMongoDBQueries, design, Node.js
MongoDB
Cheatsheet
Fifteen modules on the document database: queries and updates, schema design, aggregation, indexes and explain, then the Node.js driver and Mongoose in NestJS. Query results are computed from the demo data shown.
Key terms in plain words
New to databases? Start here. Every word below shows up later on this page, and each one comes with an everyday comparison. Underlined words in the modules link back to these cards.
34 terms in 6 groups. Hover an underlined word anywhere on the page for a quick reminder, or click it to jump here.
Where your data lives
MongoDB works like a library. The library holds shelves, the shelves hold books, and every book has its own catalogue number.
- Database
- Think of it as a whole library building
- The top level container. One MongoDB server can hold many databases, usually one per app, and each one keeps its own collections apart from the others.
- For example A shop might keep everything in a database called
shop. - Collection
- Think of it as one shelf in that library
- A group of documents of the same kind. Customers sit in one collection, orders in another. You don't have to create it first: it appears the first time you save something into it.
- For example
customers,productsandordersare three collections. - Document
- Think of it as one book on the shelf
- A single record, such as one customer or one order. It's a set of labelled values, a bit like a filled-in form, and it can hold lists and smaller forms inside it.
- For example
{ name: "Asha", city: "Pune" }is a document with two fields. - Field
- Think of it as one line on a form, like "Name" or "Phone"
- A label and the value written next to it inside a document. Two documents in the same collection can have different fields, which is part of what makes MongoDB flexible.
- For example In
{ price: 499 }the field ispriceand its value is 499. - _id and ObjectId
- Think of it as the catalogue number stamped inside every book
- Every document has an
_idthat no other document in its collection shares. If you don't give one, MongoDB makes an ObjectId for you, a 12 byte code that also records when it was created. - BSON
- Think of it as the way the library binds its books so they last
- The format MongoDB uses to store documents on disk. It looks like JSON to you, but it also knows about real dates, exact money amounts and very large whole numbers, which plain JSON cannot tell apart.
- mongosh
- Think of it as the help desk where you type requests to the librarian
- The MongoDB Shell, a program you open in a terminal to talk to the database directly. You type a command, it runs it and prints the answer.
Finding and changing things
How you ask for books, pick the pages you want, and write corrections in them.
- Query and filter
- Think of it as a request slip you hand to the librarian
- A query asks the database for documents. The filter is the part that says which ones, written as a small document of conditions. An empty filter
{}means "all of them". - For example
{ city: "Pune" }finds every customer whose city is Pune. - Operator
- Think of it as a special word on the request slip, such as "more than" or "one of"
- A keyword starting with a dollar sign that changes what a condition means.
$gtmeans greater than,$inmeans any of these values,$setmeans change this field. - Projection
- Think of it as photocopying only the pages you need instead of borrowing the whole book
- A list of the fields you want back. Leaving the rest out means less data travels over the network and your app uses less memory.
- Dot notation
- Think of it as saying "chapter 3, section 2" to point inside a book
- A way to reach a field that sits inside another field, by joining the names with dots.
- For example
"address.city"reaches the city inside a customer's address. - Upsert
- Think of it as "update my library card, or make me a new one if I don't have one"
- Update plus insert. If a matching document exists it gets changed; if not, a new one is created. One call covers both cases.
- Atomic
- Think of it as a stamp that either lands fully or not at all
- An atomic change happens completely or not at all, never half way. Two people adding to the same counter at the same moment will never lose one of the additions.
- Cursor
- Think of it as a bookmark that walks through a long reading list
- When a query matches many documents, MongoDB doesn't send them all at once. It gives you a cursor, and you pull the results through it in batches as you need them.
Planning the shape of your data
Deciding what goes in one book and what gets its own book with a note pointing to it.
- Schema
- Think of it as the blank form every new book has to follow
- The planned layout of your documents: which fields they have and what kind of value goes in each. MongoDB doesn't force one by default, but you can add rules when you want them.
- Embedding
- Think of it as printing the appendix inside the book itself
- Storing related data inside the same document. An order keeps its items and delivery address inside it, so one read gets everything.
- Reference
- Think of it as a note saying "see the book on shelf C, number 42"
- Storing only the
_idof another document instead of a copy of it. Good for data shared by many documents, such as the customer behind many orders. - Schema validation
- Think of it as a librarian at the door who checks each new book before it goes on the shelf
- Rules attached to a collection that reject documents with missing or wrongly typed fields. Every app that writes to the collection has to pass the same check.
Adding things up
Turning thousands of documents into one answer, like a sales total per city.
- Aggregation pipeline
- Think of it as a factory line where each worker does one job and passes the tray along
- A list of steps that documents flow through. One step filters, the next groups, the next sorts. What comes out the end is a report rather than raw records.
- Stage
- Think of it as one worker on that factory line
- A single step in a pipeline, written as an operator like
$match(keep some),$group(add up) or$sort(put in order). - $lookup (a join)
- Think of it as fetching the book a note points to and clipping it to the first one
- Pulls matching documents from another collection into the current one, the way an order can pull in its customer's details for a report.
Making it fast
How MongoDB avoids reading every book to find one.
- Index
- Think of it as the card catalogue at the front of the library
- A sorted list of one or more fields that points straight to the matching documents. Without one, MongoDB has to look at every document in the collection to answer a query.
- Compound index
- Think of it as a catalogue sorted by author, then by year within each author
- An index on several fields at once. The order of the fields matters: put the ones you match exactly first, then the ones you sort by, then ranges like "after this date".
- Collection scan vs index scan
- Think of it as walking every shelf compared with checking the catalogue first
COLLSCANmeans MongoDB read every document to find the answer, which gets slow as data grows.IXSCANmeans it used an index and went straight to the right ones.- Explain plan
- Think of it as asking the librarian to show their working
- A report on how MongoDB answered a query: which index it used and how many documents it had to read compared with how many it returned.
- TTL
- Think of it as a "return by" date after which the book leaves the shelf on its own
- Time to live. A TTL index deletes documents automatically once a date field is older than a set number of seconds. Handy for login sessions and one time codes.
Running it for real
Words you meet once an app depends on the database every day.
- Driver
- Think of it as a translator between your app and the librarian
- The library of code your program uses to talk to MongoDB. In Node.js it's the
mongodbpackage, andMongoClientis the object that holds the connection. - Transaction
- Think of it as a checkout where you either borrow all three books or none of them
- A group of changes across several documents that succeed together or are all undone together, so the data never ends up half changed.
- Replica set
- Think of it as a main library with branch libraries that keep exact copies
- A group of MongoDB servers holding the same data. One takes writes and the others copy it, so if one machine fails another takes over. Transactions need one, even a single server one in development.
- Write concern
- Think of it as asking for a receipt only once the branches also have a copy
- How many servers must confirm a write before MongoDB tells your app it worked. Asking for more confirmations is safer and a little slower.
- Change stream
- Think of it as a noticeboard that posts every book added, edited or removed
- A live feed of changes to a collection that your app can listen to, for example to send an email whenever a new order arrives.
- Resume token
- Think of it as a bookmark in the noticeboard's log
- A marker on every change stream event. Save the last one you handled, and after a restart you can carry on from exactly that point without missing anything.
- Mongoose and models
- Think of it as a stricter librarian who fills in blanks and checks every form for you
- Mongoose is a popular add-on for Node.js that sits on top of the driver. You describe a schema once, and the model it creates gives you defaults, validation and handy methods for reading and writing.
- Backup and restore
- Think of it as photocopying the whole library and keeping the copy somewhere else
mongodumpsaves a copy of your data to files andmongorestoreloads it back. Without a tested backup, one bad delete can be permanent.
mongosh and documents
Connect with mongosh, switch databases, and insert and read your first documents. Collections and databases appear on first write.
A document is a BSON object up to 16 MB, and every document has an _id; if you leave it out, the driver creates an ObjectId, which embeds its creation time. The demo database has four customers, four products and eight orders, and every query result on this page was computed from that data.
shop> show collections customers orders products shop> db.orders.countDocuments() 8 shop> db.orders.countDocuments({ status: 'paid' }) 3 shop> db.orders.findOne({ status: 'shipped' }) { _id: ObjectId('66f1a2b3c4d5e6f708190012'), customerId: ObjectId('66f1a2b3c4d5e6f708190002'), status: 'shipped', items: [ { sku: 'SKU-4', qty: 1, price: 4499 } ], total: 4499, shipping: { city: 'Mumbai', pincode: '400001' }, createdAt: ISODate('2026-07-06T10:30:00.000Z') } shop> db.customers.insertOne({ name: 'Esha Nair', email: 'esha@example.com', city: 'Gurugram', tier: 'bronze', createdAt: ISODate('2026-10-01T09:00:00.000Z') }) { acknowledged: true, insertedId: ObjectId('66f1a2b3c4d5e6f708190063') }
| Command | Does |
|---|---|
mongosh "mongodb://localhost:27017/shop" | Connect to a database |
show dbs / use shop | List and switch databases |
show collections | Collections in the current database |
db.orders.findOne(filter) | The first matching document |
db.orders.countDocuments(filter) | An exact count; estimatedDocumentCount() is a fast estimate |
db.orders.stats() | Size, document count and indexes |
it | Show the next batch of a cursor |
BSON types and document shape
BSON adds the types JSON lacks: ObjectId, real dates, 64 bit integers and exact decimals. Pick the right one at write time; queries depend on it.
{
_id: ObjectId('66f1a2b3c4d5e6f708190011'), // 12 bytes; the first 4 are a timestamp
customerId: ObjectId('66f1a2b3c4d5e6f708190001'), // a reference to customers._id
status: 'paid',
items: [ // an embedded array of sub documents
{ sku: 'SKU-1', qty: 1, price: 1999 },
{ sku: 'SKU-2', qty: 1, price: 1299 }
],
total: 3298, // store money as integer paise or Decimal128
shipping: { city: 'Delhi', pincode: '110001' },// embedded: read with the order, always
createdAt: ISODate('2026-07-03T10:30:00.000Z') // a real Date, so ranges and $dateToString work
}Why it matters: 1999 and "1999" are different values to MongoDB. A string stored where a number is expected never matches $gte; add schema validation (module 10) to stop that at the door.
| Type | mongosh | Use for |
|---|---|---|
| ObjectId | ObjectId() | Default _id; sortable by creation time |
| Date | ISODate('2026-10-01') / new Date() | Every timestamp; never a string |
| Int32 / Long | NumberInt(5) / NumberLong(5) | Counters and ids that must stay integers |
| Double | 3.14 | The default number type in mongosh |
| Decimal128 | NumberDecimal('19.99') | Money with fractions |
| Array | ['a', 'b'] | Lists you query by element |
| Embedded document | { city: 'Delhi' } | Data read together with its parent |
Find and query operators
Filters are documents. Combine comparison, logical, array and element operators, reach into nested fields with dot notation, and project only what you need.
A field set to a value means equals. For an array field, { tags: 'audio' } matches if any element equals it. When several conditions must hold for the same array element, wrap them in $elemMatch; without it, each condition may match a different element.
shop> db.orders.find({ status: { $in: ['paid', 'shipped'] }, total: { $gte: 3000 } }, { _id: 0, status: 1, total: 1 }).sort({ total: -1 }) [ { status: 'shipped', total: 8998 }, { status: 'paid', total: 6498 }, { status: 'shipped', total: 4499 }, { status: 'paid', total: 3897 }, { status: 'paid', total: 3298 } ] shop> db.orders.find({ 'shipping.city': 'Delhi', createdAt: { $gte: ISODate('2026-08-01T00:00:00.000Z') } }, { _id: 0, total: 1, createdAt: 1 }) [ { total: 8998, createdAt: ISODate('2026-08-17T10:30:00.000Z') }, { total: 3897, createdAt: ISODate('2026-08-22T10:30:00.000Z') } ] shop> db.orders.find({ items: { $elemMatch: { sku: 'SKU-1', qty: { $gte: 2 } } } }, { _id: 0, status: 1, items: 1 }) [ { status: 'pending', items: [ { sku: 'SKU-1', qty: 2, price: 1999 } ] } ] shop> db.products.find({ $or: [{ stock: 0 }, { tags: 'audio' }] }, { name: 1, stock: 1, tags: 1 }) [ { _id: 'SKU-1', name: 'Earbuds', stock: 40, tags: [ 'audio', 'wireless' ] }, { _id: 'SKU-3', name: 'Yoga mat', stock: 0, tags: [ 'fitness' ] }, { _id: 'SKU-4', name: 'Speaker', stock: 7, tags: [ 'audio' ] } ] shop> db.customers.find({ name: { $regex: '^[A-C]' }, tier: { $ne: 'gold' } }, { _id: 0, name: 1, tier: 1 }) [ { name: 'Ben Thomas', tier: 'silver' }, { name: 'Chen Li', tier: 'silver' } ] shop> db.products.find({ 'tags.1': { $exists: true } }, { _id: 1, tags: 1 }) [ { _id: 'SKU-1', tags: [ 'audio', 'wireless' ] } ]
Why it matters: projections cut network and memory. Returned fields keep the document's own order, and _id is included unless you set it to 0.
| Operator | Matches |
|---|---|
$eq $ne $gt $gte $lt $lte | Comparisons |
$in / $nin | Any value in a list / none of them |
$and $or $nor $not | Logical combinations |
$exists | Field is present |
$type | Field has a BSON type |
$regex | Pattern match; anchored prefixes can use an index |
$elemMatch | One array element meets every condition |
$all / $size | Array has every value / exactly n elements |
$expr | Compare two fields of the same document |
Updates and upserts
Change documents in place with update operators, upsert in one call, edit array elements by condition, and get the new document back.
Each single document write is atomic, so $inc on a counter never loses an update. findOneAndUpdate returns the document before the change unless you ask for returnDocument: 'after'. An upsert inserts when nothing matches, using $setOnInsert for fields that only belong to new documents.
shop> db.orders.updateOne({ _id: ObjectId('66f1a2b3c4d5e6f708190014') }, { $set: { status: 'paid' }, $push: { history: { status: 'paid', at: ISODate('2026-10-01T09:15:00.000Z') } } }) { acknowledged: true, insertedId: null, matchedCount: 1, modifiedCount: 1, upsertedCount: 0 } shop> db.orders.findOne({ _id: ObjectId('66f1a2b3c4d5e6f708190014') }, { _id: 0, status: 1, history: 1 }) { status: 'paid', history: [ { status: 'paid', at: ISODate('2026-10-01T09:15:00.000Z') } ] } shop> db.products.findOneAndUpdate({ _id: 'SKU-1' }, { $inc: { stock: -1 }, $addToSet: { tags: 'bestseller' } }, { returnDocument: 'after' }) { _id: 'SKU-1', name: 'Earbuds', price: 1999, stock: 39, tags: [ 'audio', 'wireless', 'bestseller' ] } shop> db.orders.updateOne({ _id: ObjectId('66f1a2b3c4d5e6f708190011') }, { $set: { 'items.$[line].qty': 2 } }, { arrayFilters: [{ 'line.sku': 'SKU-2' }] }) { acknowledged: true, insertedId: null, matchedCount: 1, modifiedCount: 1, upsertedCount: 0 } shop> db.orders.findOne({ _id: ObjectId('66f1a2b3c4d5e6f708190011') }, { _id: 0, items: 1 }) { items: [ { sku: 'SKU-1', qty: 1, price: 1999 }, { sku: 'SKU-2', qty: 2, price: 1299 } ] } shop> db.customers.updateOne({ email: 'farah@example.com' }, { $set: { city: 'Delhi' }, $setOnInsert: { name: 'Farah Khan', tier: 'bronze' } }, { upsert: true }) { acknowledged: true, insertedId: ObjectId('66f1a2b3c4d5e6f708190064'), matchedCount: 0, modifiedCount: 0, upsertedCount: 1 } shop> db.orders.updateMany({ status: 'paid', createdAt: { $lt: ISODate('2026-07-15T00:00:00.000Z') } }, { $set: { status: 'shipped' } }) { acknowledged: true, insertedId: null, matchedCount: 1, modifiedCount: 1, upsertedCount: 0 } shop> db.products.findOneAndUpdate({ _id: 'SKU-3' }, { $pull: { tags: 'fitness' }, $unset: { stock: '' } }, { returnDocument: 'after' }) { _id: 'SKU-3', name: 'Yoga mat', price: 899, tags: [] }
| Operator | Does |
|---|---|
$set / $unset | Set or remove fields |
$inc / $mul | Add to or multiply a number |
$min / $max | Update only if lower or higher |
$push with $each, $slice | Append, optionally keeping the last n |
$addToSet | Append only if absent |
$pull | Remove matching elements |
$[name] + arrayFilters | Update elements that match a condition |
$setOnInsert | Only applied when an upsert inserts |
$currentDate | Set a field to the server's time |
Schema design: embed or reference
Shape documents around how the app reads them. Embed what is read together and bounded; reference what is shared, huge or changes on its own.
The rule of thumb is that data accessed together should be stored together. An order embeds its line items and shipping address because they belong to that order and never grow without limit. It references the customer by customerId because customers are shared across many orders and change independently.
| Relationship | Usually | Example |
|---|---|---|
| One to few | Embed | Order and its line items |
| One to many, bounded | Embed or reference | Product and its 20 variants |
| One to very many | Reference from the many side | Customer and their orders |
| Many to many | Arrays of ids on one or both sides | Students and courses |
| Data that grows forever | Never an array in one document | Events, logs, comments |
| Pattern | Idea |
|---|---|
| Extended reference | Copy the few fields you always show, like the customer's name, next to the id |
| Subset | Embed the latest 10 reviews, keep the rest in their own collection |
| Bucket | Group time series points into one document per hour or day |
| Computed | Store totals on write, like total on an order, instead of summing on every read |
| Schema versioning | Add a schemaVersion field and migrate documents lazily |
Aggregation pipeline
A pipeline is a list of stages; each one takes documents in and passes documents on. Filter early, group, reshape and sort.
Put $match and $sort first so they can use indexes, then $group to summarise. Field paths start with $, and $group needs an _id: the value you group by, or null for one total.
shop> db.orders.aggregate([ ... { $match: { status: { $ne: 'cancelled' } } }, ... { $group: { _id: '$status', orders: { $sum: 1 }, revenue: { $sum: '$total' } } }, ... { $sort: { revenue: -1 } } ... ]) [ { _id: 'shipped', orders: 4, revenue: 18593 }, { _id: 'paid', orders: 3, revenue: 14393 } ] shop> db.orders.aggregate([ ... { $match: { status: { $in: ['paid', 'shipped'] } } }, ... { $group: { _id: { $dateToString: { format: '%Y-%m', date: '$createdAt' } }, revenue: { $sum: '$total' }, avgOrder: { $avg: '$total' } } }, ... { $project: { _id: 0, month: '$_id', revenue: '$revenue', avgOrder: { $round: ['$avgOrder', 0] } } }, ... { $sort: { month: 1 } } ... ]) [ { month: '2026-07', revenue: 13593, avgOrder: 3398 }, { month: '2026-08', revenue: 19393, avgOrder: 6464 } ]
$lookup, $unwind and $facet
Join collections with $lookup, flatten arrays with $unwind, and compute several summaries in one pass with $facet.
$lookup adds an array of matching documents from another collection; index the foreignField. If you need $lookup on most reads, that is a hint to embed or copy a few fields instead (module 05).
shop> db.orders.aggregate([ ... { $unwind: '$items' }, ... { $group: { _id: '$items.sku', unitsSold: { $sum: '$items.qty' } } }, ... { $lookup: { from: 'products', localField: '_id', foreignField: '_id', as: 'product' } }, ... { $project: { _id: 0, sku: '$_id', name: { $first: '$product.name' }, unitsSold: '$unitsSold' } }, ... { $sort: { unitsSold: -1, sku: 1 } }, ... { $limit: 3 } ... ]) [ { sku: 'SKU-2', name: 'Kettle', unitsSold: 6 }, { sku: 'SKU-1', name: 'Earbuds', unitsSold: 4 }, { sku: 'SKU-4', name: 'Speaker', unitsSold: 4 } ] shop> db.customers.aggregate([ ... { $match: { city: 'Delhi' } }, ... { $lookup: { from: 'orders', localField: '_id', foreignField: 'customerId', as: 'orders' } }, ... { $project: { _id: 0, name: '$name', orderCount: { $size: '$orders' }, spent: { $sum: '$orders.total' } } }, ... { $sort: { spent: -1 } } ... ]) [ { name: 'Asha Rao', orderCount: 3, spent: 14094 }, { name: 'Chen Li', orderCount: 2, spent: 7895 }, { name: 'Farah Khan', orderCount: 0, spent: 0 } ] shop> db.orders.aggregate([ ... { $facet: { byStatus: [{ $sortByCount: '$status' }], bySize: [{ $bucket: { groupBy: '$total', boundaries: [0, 2000, 5000, 100000], default: 'other', output: { count: { $sum: 1 } } } }] } } ... ]) [ { byStatus: [ { _id: 'shipped', count: 4 }, { _id: 'paid', count: 3 }, { _id: 'cancelled', count: 1 } ], bySize: [ { _id: 0, count: 2 }, { _id: 2000, count: 4 }, { _id: 5000, count: 2 } ] } ]
Why it matters: $facet runs sub pipelines over the same input, which is how a search page gets results and filter counts in a single query.
Indexes
Single field, compound, multikey, unique, partial, TTL and text indexes, and the ESR rule for ordering compound index keys.
Order compound keys by the ESR rule: fields matched by Equality first, then Sort fields, then Range fields. A query on { customerId, createdAt > x } sorted by date is served by { customerId: 1, createdAt: -1 }. An index on an array field is multikey: one entry per element.
shop> db.orders.createIndex({ customerId: 1, createdAt: -1 }) customerId_1_createdAt_-1 shop> db.customers.createIndex({ email: 1 }, { unique: true }) email_1 shop> db.orders.createIndex({ createdAt: 1 }, { partialFilterExpression: { status: 'pending' } }) createdAt_1 shop> db.sessions.createIndex({ expiresAt: 1 }, { expireAfterSeconds: 0 }) expiresAt_1 shop> db.products.createIndex({ name: 'text', tags: 'text' }) name_text_tags_text shop> db.orders.getIndexes().map(i => i.name) [ '_id_', 'customerId_1_createdAt_-1', 'createdAt_1' ]
| Index | Good for | Note |
|---|---|---|
| Compound | Filters and sorts on several fields | Prefixes are usable alone; order matters |
| Multikey | Queries on array elements | Automatic when the field holds arrays |
| Unique | Enforcing one per value | Combine with partial for optional fields |
| Partial | A hot subset, such as pending orders | Query must include the filter condition |
| TTL | Expiring sessions, tokens, logs | A background task deletes about once a minute |
| Text | Simple keyword search | Atlas Search is the fuller option |
| Wildcard | Querying unknown keys in a sub document | { 'attrs.$**': 1 } |
Reading explain()
explain('executionStats') shows the winning plan and how much work it did. Compare documents examined with documents returned.
Look for IXSCAN rather than COLLSCAN, no in memory SORT stage, and totalDocsExamined close to nReturned. Without the compound index the same query shows a COLLSCAN over every order followed by a SORT.
shop> db.orders.find({ customerId: ObjectId('66f1a2b3c4d5e6f708190001') }) ... .sort({ createdAt: -1 }).limit(5).explain('executionStats') { queryPlanner: { namespace: 'shop.orders', winningPlan: { stage: 'LIMIT', limitAmount: 5, inputStage: { stage: 'FETCH', inputStage: { stage: 'IXSCAN', keyPattern: { customerId: 1, createdAt: -1 }, indexName: 'customerId_1_createdAt_-1', direction: 'forward', indexBounds: { customerId: [ "[ObjectId('66f1a2b3c4d5e6f708190001'), ObjectId('66f1a2b3c4d5e6f708190001')]" ], createdAt: [ '[MaxKey, MinKey]' ] } } } }, rejectedPlans: [] }, executionStats: { executionSuccess: true, nReturned: 3, executionTimeMillis: 0, totalKeysExamined: 3, totalDocsExamined: 3, ... }, ... }
Why it matters: three keys examined, three documents read, three returned: the index found exactly the right rows, already in order. On a real collection, a ratio far above 1 means the index does not fit the query.
| Stage | Means |
|---|---|
COLLSCAN | Read every document |
IXSCAN | Walked an index |
FETCH | Loaded documents found by the index |
SORT | Sorted in memory; capped at 100 MB, so index the sort |
PROJECTION_COVERED | Answered from the index alone |
SHARDING_FILTER | Dropped documents owned by another shard |
Schema validation
MongoDB is schema flexible, not schemaless. A $jsonSchema validator rejects documents that do not match, for every client that writes.
db.createCollection('orders', {
validator: {
$jsonSchema: {
bsonType: 'object',
required: ['customerId', 'status', 'items', 'total', 'createdAt'],
properties: {
customerId: { bsonType: 'objectId' },
status: { enum: ['pending', 'paid', 'shipped', 'cancelled'] },
items: {
bsonType: 'array',
minItems: 1,
items: {
bsonType: 'object',
required: ['sku', 'qty', 'price'],
properties: { qty: { bsonType: 'int', minimum: 1 }, price: { bsonType: ['int', 'long', 'decimal'] } }
}
},
total: { bsonType: ['int', 'long', 'decimal'], minimum: 0 },
createdAt: { bsonType: 'date' }
}
}
},
validationLevel: 'strict', // or 'moderate': skip documents that were already invalid
validationAction: 'error' // or 'warn': log instead of rejecting
});
// Change the rules on an existing collection
db.runCommand({ collMod: 'orders', validator: { /* ... */ }, validationLevel: 'moderate' });shop> db.orders.insertOne({ customerId: 'c1', status: 'new', items: [], total: -5 }) MongoServerError: Document failed validation Additional information: { failingDocumentId: ObjectId('...'), details: { operatorName: '$jsonSchema', schemaRulesNotSatisfied: [ ... ] } }
From Node.js with the driver
Typed collections, projections, cursors, bulk writes, aggregation, transactions and change streams with the official driver.
Create one MongoClient per process and reuse it; it holds the connection pool. Passing a document type to collection<Order>() makes filters and updates type checked. This file compiles in strict mode against the mongodb 7.7 driver.
import { MongoClient, ObjectId, type Collection } from 'mongodb';
interface LineItem { sku: string; qty: number; price: number }
interface Order {
_id: ObjectId;
customerId: ObjectId;
status: 'pending' | 'paid' | 'shipped' | 'cancelled';
items: LineItem[];
total: number;
createdAt: Date;
}
// One client per process; it manages its own connection pool
const client = new MongoClient(process.env.MONGO_URL ?? 'mongodb://localhost:27017/?replicaSet=rs0', {
maxPoolSize: 20,
retryWrites: true,
});
await client.connect();
const db = client.db('shop');
const orders: Collection<Order> = db.collection<Order>('orders');
// Indexes are idempotent: run them at startup or in a migration
await orders.createIndex({ customerId: 1, createdAt: -1 });
await orders.createIndex({ status: 1, createdAt: 1 }, { partialFilterExpression: { status: 'pending' } });
// Typed queries: filters, projections and results all checked against Order
const recent = await orders
.find({ customerId: new ObjectId('66f1a2b3c4d5e6f708190001'), status: { $ne: 'cancelled' } })
.project<Pick<Order, 'total' | 'createdAt'>>({ _id: 0, total: 1, createdAt: 1 })
.sort({ createdAt: -1 })
.limit(5)
.toArray();
// Cursors stream large results instead of loading them all
for await (const order of orders.find({ status: 'paid' }).batchSize(500)) {
void order.total;
}
// bulkWrite: many different writes, one round trip
await orders.bulkWrite([
{ updateOne: { filter: { status: 'paid', createdAt: { $lt: new Date('2026-07-15') } }, update: { $set: { status: 'shipped' } } } },
{ deleteMany: { filter: { status: 'cancelled', createdAt: { $lt: new Date('2025-01-01') } } } },
], { ordered: false });
// Typed aggregation output
const revenue = await orders.aggregate<{ _id: string; revenue: number }>([
{ $match: { status: { $in: ['paid', 'shipped'] } } },
{ $group: { _id: '$status', revenue: { $sum: '$total' } } },
]).toArray();
// A multi document transaction (needs a replica set or a sharded cluster)
const session = client.startSession();
try {
await session.withTransaction(async () => {
const stock = db.collection<{ _id: string; stock: number }>('products');
const res = await stock.updateOne({ _id: 'SKU-1', stock: { $gte: 1 } }, { $inc: { stock: -1 } }, { session });
if (res.modifiedCount === 0) throw new Error('Out of stock'); // aborts the transaction
await orders.insertOne({
_id: new ObjectId(), customerId: new ObjectId(), status: 'pending',
items: [{ sku: 'SKU-1', qty: 1, price: 1999 }], total: 1999, createdAt: new Date(),
}, { session });
}, { readConcern: { level: 'snapshot' }, writeConcern: { w: 'majority' } });
} finally {
await session.endSession();
}
// Change streams: react to writes as they happen, and resume after a restart
const stream = orders.watch([{ $match: { operationType: 'insert', 'fullDocument.total': { $gte: 5000 } } }]);
stream.on('change', (event) => {
if (event.operationType === 'insert') console.log('big order', event.fullDocument._id, event._id);
});
console.log(recent, revenue);
await stream.close();
await client.close();Transactions and consistency
Single document writes are already atomic. Use multi document transactions when several documents must change together, and choose read and write concerns.
Transactions need a replica set or sharded cluster, even a one node replica set in development. withTransaction retries on transient errors for you. Keep them short: under a second, and touching as few documents as you can. Often a better document shape removes the need for one.
| Setting | Means | Use |
|---|---|---|
w: 1 | Acknowledged by the primary | Fast, can be rolled back on failover |
w: 'majority' | Acknowledged by most members | The default; survives failover |
readConcern 'local' | Latest data on that node | Default reads |
readConcern 'majority' | Data that cannot be rolled back | Reads that must not see lost writes |
readConcern 'snapshot' | One point in time | Inside transactions |
readPreference | Which member serves reads | secondaryPreferred for reports |
Change streams
Subscribe to inserts, updates and deletes on a collection, database or cluster, filter them with a pipeline, and resume after a restart.
Every event carries a resume token in _id. Store the last one you processed and pass it as resumeAfter when your service restarts, and you continue exactly where you left off. Updates include only the changed fields unless you ask for fullDocument: 'updateLookup'.
const stream = orders.watch(
[{ $match: { operationType: { $in: ['insert', 'update'] }, 'fullDocument.status': 'paid' } }],
{ fullDocument: 'updateLookup', resumeAfter: await loadLastToken() },
);
for await (const event of stream) {
if (event.operationType === 'insert' || event.operationType === 'update') {
await sendReceipt(event.fullDocument!);
}
await saveLastToken(event._id); // the resume token
}Mongoose in NestJS
Define schemas with decorators, register models per module, and inject them into services with @InjectModel.
Mongoose adds schemas, validation, defaults and middleware on top of the driver. lean() skips building full Mongoose documents and returns plain objects, which is much faster for reads you only serialise. These three files compile against NestJS 12, @nestjs/mongoose 12 and Mongoose 9.
import { Prop, Schema, SchemaFactory } from '@nestjs/mongoose';
import { HydratedDocument, Types } from 'mongoose';
@Schema({ _id: false })
export class LineItem {
@Prop({ required: true }) sku!: string;
@Prop({ required: true, min: 1 }) qty!: number;
@Prop({ required: true, min: 0 }) price!: number;
}
@Schema({ collection: 'orders', timestamps: true }) // adds createdAt and updatedAt
export class Order {
@Prop({ type: Types.ObjectId, ref: 'Customer', required: true, index: true })
customerId!: Types.ObjectId;
@Prop({ enum: ['pending', 'paid', 'shipped', 'cancelled'], default: 'pending' })
status!: string;
@Prop({ type: [SchemaFactory.createForClass(LineItem)], default: [] })
items!: LineItem[];
@Prop({ required: true, min: 0 })
total!: number;
}
export type OrderDocument = HydratedDocument<Order>;
export const OrderSchema = SchemaFactory.createForClass(Order);
OrderSchema.index({ customerId: 1, createdAt: -1 });import { Module } from '@nestjs/common';
import { MongooseModule } from '@nestjs/mongoose';
import { ConfigService } from '@nestjs/config';
import { Order, OrderSchema } from './order.schema.js';
import { OrdersService } from './orders.service.js';
@Module({
imports: [
MongooseModule.forRootAsync({
inject: [ConfigService],
useFactory: (config: ConfigService) => ({ uri: config.getOrThrow<string>('MONGO_URL') }),
}),
MongooseModule.forFeature([{ name: Order.name, schema: OrderSchema }]),
],
providers: [OrdersService],
exports: [OrdersService],
})
export class OrdersModule {}import { Injectable, NotFoundException } from '@nestjs/common';
import { InjectModel } from '@nestjs/mongoose';
import { Model, Types } from 'mongoose';
import { Order } from './order.schema.js';
@Injectable()
export class OrdersService {
constructor(@InjectModel(Order.name) private readonly orders: Model<Order>) {}
// lean() returns plain objects: faster, no document methods
recentFor(customerId: string) {
return this.orders
.find({ customerId: new Types.ObjectId(customerId) })
.sort({ createdAt: -1 })
.limit(5)
.select({ total: 1, status: 1, createdAt: 1 })
.lean();
}
async markPaid(id: string) {
const order = await this.orders.findByIdAndUpdate(id, { $set: { status: 'paid' } }, { new: true, runValidators: true });
if (!order) throw new NotFoundException(`Order ${id} not found`);
return order;
}
revenueByStatus() {
return this.orders.aggregate<{ _id: string; revenue: number }>([
{ $group: { _id: '$status', revenue: { $sum: '$total' } } },
{ $sort: { revenue: -1 } },
]);
}
}Why it matters: runValidators: true makes update queries check the schema too. Mongoose runs validators on save by default but not on updates.
Operations
Back up and restore, find slow operations with the profiler, kill a runaway query, and check sizes and replica set health.
# Back up one database to a compressed archive, then restore it elsewhere
mongodump --uri="mongodb://localhost:27017/shop" --gzip --archive=shop.archive.gz
mongorestore --uri="mongodb://localhost:27017" --gzip --archive=shop.archive.gz --nsFrom='shop.*' --nsTo='shop_copy.*'
# Export one collection as JSON lines, import it back
mongoexport --uri="mongodb://localhost:27017/shop" --collection=orders --out=orders.jsonl
mongoimport --uri="mongodb://localhost:27017/shop" --collection=orders --file=orders.jsonl// Log operations slower than 100 ms into system.profile
db.setProfilingLevel(1, { slowms: 100 })
db.system.profile.find().sort({ ts: -1 }).limit(5)
// Operations running longer than 5 seconds, then stop one
db.currentOp({ active: true, secs_running: { $gt: 5 } })
db.killOp(12345)
// Sizes, in MB
db.stats(1024 * 1024)
db.orders.totalIndexSize() / 1024 / 1024
// Replica set health and replication lag
rs.status().members.map(m => ({ name: m.name, state: m.stateStr, lagSecs: m.optimeDate }))
rs.printSecondaryReplicationInfo()| Watch | Why |
|---|---|
| Working set vs RAM | Indexes and hot documents should fit in the WiredTiger cache |
| Slow query log | Queries with COLLSCAN or a high docs examined ratio |
| Replication lag | Secondaries far behind risk stale reads and slow failover |
| Connections | A climbing count usually means clients are not reusing MongoClient |
| Oplog window | Must be longer than your longest maintenance or outage |
Which one do I need?
Start from what you are trying to do, then reach for the tool in the middle column. The last column takes you to the module that explains it.
| I want to | Reach for | Example | Module |
|---|---|---|---|
| Read one document | findOne | db.orders.findOne({ _id }) | 01 |
| Store money safely | Integer minor units or Decimal128 | NumberDecimal('19.99') | 02 |
| Match an array element on several fields | $elemMatch | { items: { $elemMatch: {...} } } | 03 |
| Return only some fields | Projection | find(f, { _id: 0, total: 1 }) | 03 |
| Increment a counter atomically | $inc | { $inc: { stock: -1 } } | 04 |
| Insert if missing | Upsert | updateOne(f, u, { upsert: true }) | 04 |
| Edit matching array elements | arrayFilters | 'items.$[line].qty' | 04 |
| Decide what to embed | Read patterns | Embed bounded, reference shared | 05 |
| Totals per group | $group | { $group: { _id: '$status', n: { $sum: 1 } } } | 06 |
| Join another collection | $lookup | { $lookup: { from, localField, ... } } | 07 |
| Several summaries at once | $facet | { $facet: { a: [...], b: [...] } } | 07 |
| Speed up filter and sort | Compound index, ESR order | { customerId: 1, createdAt: -1 } | 08 |
| Expire documents automatically | TTL index | expireAfterSeconds: 0 | 08 |
| Check a query uses an index | explain | .explain('executionStats') | 09 |
| Reject malformed writes | $jsonSchema validator | validator: { $jsonSchema } | 10 |
| Type check queries in Node | Collection | db.collection<Order>('orders') | 11 |
| Change several documents together | Transaction | session.withTransaction() | 12 |
| React to writes | Change stream | orders.watch(pipeline) | 13 |
| Models in NestJS | @nestjs/mongoose | @InjectModel(Order.name) | 14 |
| Back up a database | mongodump | mongodump --gzip --archive | 15 |