Explainer cheatsheetFrom one electron to the whole driveSimple, analogy, technical
SSD
Unpacked
An SSD looks like a stick of gum with a few chips on it. Inside, billions of cells hold trapped electrons, a small computer moves your data around constantly, and a map keeps every file findable. Take one apart here, layer by layer, with each idea in simple words, an everyday picture and the full technical detail.
The life of one writeTap a box to jump to its module
Where storage fits
Storage is the slow but permanent end of the memory ladder. RAM forgets when power goes off; an SSD remembers for years.
In simple words
A computer keeps what it is working on right now in fast, small memory (RAM), and everything it must keep in slower, much bigger storage. An SSD stores data in flash memory chips with no moving parts. A hard disk stores it on spinning magnetic platters. Both are non-volatile: they keep data with the power off.
Think of it like
Your desk and your cupboard. Papers you are using sit on the desk (RAM): quick to grab, little space, cleared every evening. Everything else lives in the cupboard (storage): more space, a few steps away, still there tomorrow. An SSD is a cupboard right next to the desk; a hard disk is one down the corridor.
Swipe sideways to see the whole diagram
- 1Each step down is roughly 10 to 1,000 times slower, which is why caching exists everywhere.
- 2An NVMe SSD is about 100 times faster than a hard disk for small random reads.
- 3The bars are on a log scale: each tick is ten times the last.
| Level | Typical access | If 1 ns were 1 second |
|---|---|---|
| L1 cache | 1 ns | 1 second |
| RAM | 100 ns | About 2 minutes |
| NVMe SSD | 60 µs | About 17 hours |
| Hard disk | 6 ms | About 2 months |
Technical deep diveTap to fold
- Volatile vs non-volatile
- DRAM stores bits as charge in capacitors that leak, so it must be refreshed constantly and loses everything without power. NAND flash traps charge behind insulators and holds it for years.
- Latency
- Time for one operation to complete. For SSDs it is usually quoted as 4 KiB random read latency at queue depth 1.
- Throughput
- Bytes per second for large sequential transfers, quoted in MB/s or GB/s.
- IOPS
- Input/output operations per second, usually 4 KiB random reads or writes. Consumer NVMe drives reach hundreds of thousands to over a million at high queue depth.
- Block device
- To the OS, any SSD or disk is a numbered array of fixed size sectors (512 B or 4 KiB) it can read and write. Everything inside is the drive's business.
An SSD, taken apart
An SSD is a small computer of its own: flash chips that hold the data, a controller that manages them, a little RAM, and power circuitry, on one circuit board.
In simple words
Lift the label off an M.2 SSD and you find a handful of chips on a board. The big black ones are NAND flash where your files live. The square one next to the connector is the controller, a processor that decides where every byte goes. A small DRAM chip holds its map, and the gold fingers at the end plug into the computer.
Think of it like
An SSD is a warehouse. The NAND chips are the shelves, the controller is the warehouse manager with a clipboard, the DRAM is that clipboard listing where every box is, and the connector is the loading dock where trucks (your computer) drop off and pick up.
Swipe sideways to see the whole diagram
- 1Label or heatsink: thin foil that also spreads heat; some drives add a metal heatsink.
- 2NAND flash packages: each holds several stacked dies with the actual memory cells.
- 3Controller: a multi-core processor running the drive's firmware.
- 4DRAM cache: stores the address map. Budget drives skip it and borrow host RAM (HMB).
- 5PMIC: power management, turning 3.3 V into the voltages each chip needs.
- 6PCB: the multi-layer circuit board wiring it all together.
- 7M.2 edge connector: gold fingers carrying PCIe lanes and power. The notch is the M key.
- 8Screw notch: holds the drive flat in the slot.
| Part | Job | Think of it as |
|---|---|---|
| NAND flash | Stores the data, even without power | Warehouse shelves |
| Controller | Maps, schedules, corrects errors, levels wear | The warehouse manager |
| DRAM cache | Holds the logical to physical map | The manager's clipboard |
| Firmware | The controller's software | The manager's rulebook |
| PMIC | Supplies clean power to every chip | The building's electrics |
| Capacitors (PLP) | Finish writes if power cuts out (enterprise drives) | A backup generator |
| Connector | Data and power to the host | The loading dock |
Technical deep diveTap to fold
- Controller
- Typically several Arm cores plus hardware engines for ECC, encryption (AES-256) and compression, with 4 to 8 channels that talk to NAND dies in parallel.
- Channels and dies
- Each channel is a bus to several dies. Speed comes from keeping every die busy at once, so a 2 TB drive with more dies is often faster than the 500 GB version of the same model.
- DRAM
- Usually about 1 GB of DRAM per 1 TB of NAND, because the mapping table needs roughly 4 bytes per 4 KiB page.
- HMB
- Host Memory Buffer: a DRAM-less NVMe drive borrows tens of MB of system RAM over PCIe to cache part of its map.
- PLP
- Power Loss Protection: capacitors that hold enough charge to flush in-flight data to NAND if power fails. Standard on data centre drives.
- Form factors
- M.2 2230, 2242 and 2280 (width 22 mm, length in mm), 2.5 inch SATA, U.2, and data centre EDSFF (E1.S, E3.S).
The memory cell
Every bit lives in a transistor that can trap electrons in an insulated layer. Trapped electrons stay put for years, even unpowered.
In simple words
A flash cell is a tiny switch with an extra pocket inside. Writing pushes electrons into the pocket; the insulation around it keeps them there. Reading checks how hard it is to turn the switch on: a full pocket makes it harder. So the amount of trapped charge is the data.
Think of it like
Think of a sealed jar you can fill with marbles through a one way funnel. To read it, you weigh the jar instead of opening it. A light jar means 1, a heavy jar means 0. Emptying the jar needs a strong shake, which is why erasing is slow and wears the jar out a little each time.
Swipe sideways to see the whole diagram
- 1The control gate is connected to a word line shared by a row of cells.
- 2Electrons sit in the charge trap layer, insulated above and below by oxide.
- 3More trapped charge raises the cell's threshold voltage, the gate voltage needed to let current flow.
Technical deep diveTap to fold
- Program
- About 20 V on the control gate pulls electrons through the thin tunnel oxide into the trap by Fowler-Nordheim tunnelling. It is done in small steps with checks in between (ISPP) to land on an exact charge level.
- Read
- The controller applies a reference voltage to the word line and senses whether the string conducts. Several references tell several charge levels apart.
- Erase
- A high voltage on the substrate pulls electrons out of every cell in a whole block at once.
- Charge trap vs floating gate
- Older planar NAND stored charge on a conductive floating gate. Almost all 3D NAND uses an insulating silicon nitride trap instead, which leaks less between neighbouring cells.
- Wear
- Each program and erase damages the tunnel oxide slightly, letting charge leak faster over time. That is the physical reason flash has limited endurance.
- Retention
- How long data survives unpowered. It shrinks as a cell wears and as temperature rises, which is why old, heavily used drives should not sit in a drawer for years.
Bits per cell: SLC to QLC
Squeezing more bits into each cell makes drives cheaper and bigger, at the cost of speed and endurance.
In simple words
A cell can hold just "empty or full" (1 bit), or several fill levels in between. Four levels store 2 bits, eight levels store 3 bits, sixteen store 4. Most drives today use TLC (3 bits per cell); cheap large ones use QLC (4 bits).
Think of it like
Reading a glass of water. If you only ask "empty or full?" anyone can answer instantly and spills do not matter. If you must tell sixteen fill levels apart, you need a careful look, it takes longer, and a little evaporation changes the answer. That is QLC.
Swipe sideways to see the whole diagram
- 1Each hump is the spread of cells programmed to one level.
- 2More levels means narrower gaps, so reads need finer reference voltages and more error correction.
- 3Programming must place charge more precisely, so writes are slower.
| Type | Bits per cell | Typical P/E cycles | Where you find it |
|---|---|---|---|
| SLC | 1 | 50,000 to 100,000 | Industrial drives, and as a fast cache inside other drives |
| MLC | 2 | 3,000 to 10,000 | Older and some enterprise drives |
| TLC | 3 | 1,000 to 3,000 | Most laptops, desktops and servers |
| QLC | 4 | Several hundred to about 1,000 | Large, cheap, read-heavy storage |
Technical deep diveTap to fold
- P/E cycles
- Program/erase cycles a block can survive before it becomes unreliable. Figures vary by generation and vendor; modern ECC stretches them.
- Read retry
- When a read fails ECC, the controller shifts its reference voltages and tries again, which adds latency on worn or hot drives.
- pSLC
- Pseudo-SLC: running TLC or QLC cells in 1 bit mode for speed and endurance. This is how the SLC cache in module 10 works.
- Cost per bit
- Going from TLC to QLC adds 33% capacity from the same silicon, which is why QLC dominates large, cheap drives.
How cells are organised
Billions of cells are grouped into pages, pages into blocks, blocks into planes and dies, and the whole thing is built upward in hundreds of layers.
In simple words
Cells are wired in rows. A row of cells read and written together is a page, usually 16 KB. Hundreds of pages form a block. Thousands of blocks sit on a die, and several dies are stacked inside each black package on the board. Modern chips build cells upward in more than 200 layers, called 3D NAND.
Think of it like
A library. A cell is one letter, a page is a page of a book, a block is a whole book, a plane is a bookcase, a die is a floor of the library and the package is the building. 3D NAND is adding floors instead of buying more land.
Swipe sideways to see the whole diagram
- 1A package contains one or more stacked dies.
- 2Each die has 2 to 6 planes that can work in parallel.
- 3A block is made of rows of pages; the highlighted row is one page.
Swipe sideways to see the whole diagram
- 1Horizontal layers are word lines; each layer is one control gate.
- 2Each violet pillar is a vertical NAND string: a column of cells in series.
- 3Adding layers adds capacity without shrinking cells, which keeps them reliable.
Technical deep diveTap to fold
- Page size
- Typically 16 KiB of data plus spare bytes for ECC and metadata. The host thinks in 4 KiB sectors, so the controller packs four into each page.
- Block size
- Hundreds to over a thousand pages, so several MB to tens of MB on modern TLC.
- NAND string
- Cells connected in series between a bit line and the source line. Reading one cell means turning all the others in the string fully on.
- Layer count
- Shipping 3D NAND is past 200 layers, with the newest above 300. Very tall stacks are made in two or more decks bonded together.
- Parallelism
- Throughput = channels × dies × planes working at once. The controller stripes writes across all of them.
Writing and erasing
Flash has one strange rule: you can write a page only if it is empty, and you can only empty a whole block at a time.
In simple words
You cannot change a page that already holds data. To update a file, the SSD writes the new version to an empty page somewhere else and marks the old page as stale. Stale pages pile up until the drive erases their whole block, which it can only do after moving any still valid pages out.
Think of it like
A notebook written in pen, where the only eraser is ripping out a whole sheet of 200 lines. To fix one line you cross it out and write the correction on the next blank line. Eventually a sheet is mostly crossings out, so you copy the few good lines to a new sheet and tear the old one out.
Swipe sideways to see the whole diagram
- 1The old page in Block A becomes stale: still physically there, but no longer referenced.
- 2The new data goes to the next free page in another block: an out of place write.
- 3Only an erase turns pages back into empty ones, and it covers the entire block.
| Operation | Unit | Typical time |
|---|---|---|
| Read | Page (16 KiB) | About 50 µs |
| Program (write) | Page | Hundreds of µs |
| Erase | Block (MBs) | A few ms |
Technical deep diveTap to fold
- Asymmetry
- Reads are fast, programs slower, erases slowest and coarse grained. Every clever trick in the controller exists to hide this.
- Sequential programming
- Pages inside a block must be programmed in order, to limit disturbance of already written neighbours.
- Read disturb and program disturb
- Reading or writing a cell nudges charge in nearby cells. The controller counts reads per block and rewrites blocks before errors build up.
- Erase suspend
- Long erases can be paused so an urgent read is served first, which keeps read latency low during background work.
The FTL: a map of everything
Your OS thinks it writes to fixed addresses. The flash translation layer quietly moves data around and keeps a map so every address still finds its data.
In simple words
The computer asks for "sector 2" and expects it to stay at sector 2 forever. Inside, the SSD keeps moving data to fresh pages. The flash translation layer is a lookup table that records, for every logical address, which physical page currently holds it. Every write updates the table.
Think of it like
Hotel room keys. You are always "the guest in booking 2". Behind the desk, staff may move your luggage to a different room while you are out, and update the register. You ask for booking 2 and still get your luggage, without ever knowing it moved.
Swipe sideways to see the whole diagram
- 1The host uses LBAs, logical block addresses, numbered 0 to the drive's size.
- 2The mapping table translates each LBA into block and page.
- 3An update writes a new page and changes one entry; the old page becomes stale.
Technical deep diveTap to fold
- Page-level mapping
- One entry per 4 KiB of data: flexible and fast, but needs about 1 GB of DRAM per TB. Almost every modern SSD uses it.
- DRAM-less drives
- Keep only part of the map in on-chip SRAM or HMB and fetch the rest from NAND, which slows random access to large files.
- Map persistence
- The table lives in DRAM for speed and is saved to NAND regularly. After a sudden power loss, the controller rebuilds recent entries from metadata stored with each page.
- Logical vs physical
- Overwriting a file at the same LBA never touches the same flash twice in a row. Tools that "securely overwrite" files on an SSD therefore do not work; use the drive's sanitize or secure erase command instead.
Garbage collection and TRIM
Stale pages must be cleared to make room. The SSD does it in the background, and TRIM tells it which pages it can forget.
In simple words
When free blocks run low, the controller picks a block that is mostly stale, copies its few valid pages elsewhere, then erases the whole block so it is empty again. This is garbage collection. When you delete a file, the OS sends TRIM so the drive knows those pages are garbage too, instead of carefully copying deleted data around.
Think of it like
Tidying the pen notebook from module 06. You pick the sheet with the most crossed out lines, copy the few good lines onto a fresh sheet, and tear the old one out. TRIM is someone telling you "that whole chapter was cancelled", so you can skip copying it.
Swipe sideways to see the whole diagram
- 1Pick a victim block with the most stale pages.
- 2Copy its valid pages into a free block.
- 3Erase the victim, which now joins the free pool.
Technical deep diveTap to fold
- Write amplification factor
WAF = bytes written to NAND ÷ bytes written by the host. 1.0 is ideal; random small writes on a full drive can push it above 3, wearing NAND faster and slowing writes.- Over-provisioning
- Hidden spare capacity, at least the ~7% difference between GB and GiB, more on enterprise drives. More spare space means emptier victim blocks and lower WAF.
- TRIM
- ATA TRIM on SATA, Deallocate (Dataset Management) on NVMe. Linux runs it weekly with
fstrim.timeron most distros, or continuously with thediscardmount option. - Background vs foreground GC
- Drives clean up while idle. A drive kept full and hammered with writes must clean while you wait, which shows up as latency spikes.
- Keep free space
- Leaving 10 to 20% of a consumer SSD empty keeps garbage collection cheap and performance steady.
Wear leveling and endurance
Each block survives a limited number of erases, so the controller spreads writes evenly and corrects errors as cells age.
In simple words
If the same few blocks were erased over and over, they would wear out while others stayed new. Wear leveling rotates writes across every block, even moving rarely changed data now and then so its blocks get a turn. Error correction fixes the bits that drift as cells age, and worn blocks are retired.
Think of it like
Rotating the tyres on a car. Front tyres wear faster, so you swap them around regularly and all four wear out together, much later than the front pair would alone.
Swipe sideways to see the whole diagram
- 1Dynamic wear leveling sends new writes to the least worn free blocks.
- 2Static wear leveling also moves cold data so its blocks join the rotation.
- 3The drive's life ends when spare blocks run out, not when the first cell fails.
Technical deep diveTap to fold
- ECC
- Modern controllers use LDPC codes that can fix many bad bits per KiB, with soft decoding using multiple reads when needed.
- Bad block management
- Factory bad blocks are mapped out at birth; blocks that fail later are retired and replaced from the spare pool.
- TBW
- Terabytes written: the vendor's write warranty. A 1 TB consumer TLC drive is often rated around 600 TBW.
- DWPD
- Drive writes per day over the warranty period, used for data centre drives. 1 DWPD on a 4 TB drive means 4 TB of writes every day for five years.
- Percentage used
- NVMe SMART reports wear as a percentage of rated life. 100% does not mean instant death, but plan replacement.
Caches: DRAM and SLC
Drives look fastest in short bursts. A fast SLC cache takes writes first; once it fills, you see the real speed of the flash.
In simple words
Writing TLC or QLC cells precisely is slow, so drives first write incoming data quickly in 1 bit per cell mode, an SLC cache. Later, when idle, they rewrite it compactly. Copy a huge file and the cache fills, and speed drops to the flash's natural pace. A DRAM chip, if present, keeps the address map fast.
Think of it like
A busy restaurant with a big walk-in fridge. Deliveries go straight into the fridge in a rush (fast), and the kitchen sorts them onto proper shelves later. If an enormous delivery fills the fridge, the trucks must wait while staff shelve things one by one.
Swipe sideways to see the whole diagram
- 1First writes land in the SLC cache at full speed.
- 2When it fills, writes go straight to TLC: the sustained write speed.
- 3Some drives drop further while folding cached data into TLC at the same time.
Technical deep diveTap to fold
- Static vs dynamic SLC cache
- Static: a fixed reserved area. Dynamic: free space used in SLC mode, so the cache shrinks as the drive fills. A nearly full drive has a tiny cache.
- Folding
- Rewriting three SLC pages into one TLC page. It runs while idle, and also during heavy writes once space is tight.
- Read caching
- NAND reads are already fast; DRAM mostly speeds up the map lookup before each read, which matters for random reads across a large drive.
- Benchmarks
- Box speeds are sequential, high queue depth, cache-sized bursts. For servers and databases, look at sustained write speed, 4 KiB random IOPS at low queue depth, and latency percentiles.
Interfaces: SATA and NVMe
The same flash can sit behind a slow old interface or a fast modern one. NVMe talks to the CPU directly over PCIe with many parallel queues.
In simple words
SATA was designed for spinning hard disks, which can only do one thing at a time, so it has one short queue and tops out around 550 MB/s. NVMe was designed for flash: every CPU core gets its own queue, and data travels over PCIe lanes straight to the processor, many times faster.
Think of it like
A single checkout counter with one line of 32 shoppers (SATA), versus a supermarket with a checkout for every aisle and room for thousands in each line (NVMe). The products on the shelves are the same; only the way out changed.
Swipe sideways to see the whole diagram
- 1SATA uses the AHCI protocol: one command queue of 32 entries (NCQ).
- 2NVMe allows up to 65,535 I/O queues with up to 65,536 commands each.
- 3Each PCIe lane adds bandwidth; M.2 NVMe drives use four.
| Interface | Typical max | Connector |
|---|---|---|
| SATA III | ≈ 550 MB/s | 2.5 inch SATA, M.2 B+M key |
| NVMe PCIe 3.0 x4 | ≈ 3.5 GB/s | M.2 M key, U.2 |
| NVMe PCIe 4.0 x4 | ≈ 7 GB/s | M.2 M key, U.2, E1.S |
| NVMe PCIe 5.0 x4 | ≈ 14 GB/s | M.2 M key, E1.S, E3.S |
Technical deep diveTap to fold
- Submission and completion queues
- The host writes commands into a ring buffer in RAM and rings a doorbell register; the drive fetches them, executes, and posts results to a completion queue, then interrupts the core that asked.
- Namespace
- An NVMe drive exposes one or more namespaces, each a separate block device such as
/dev/nvme0n1. - M.2 keys
- M key: PCIe x4, used by NVMe. B key: SATA or PCIe x2. A drive with both notches is almost always SATA. The slot must support the protocol, not just the shape.
- Linux I/O scheduler
- NVMe devices default to
none, because the drive does its own scheduling across its queues. - NVMe over Fabrics
- The same command set carried over RDMA or TCP, so remote flash in a storage server behaves almost like a local drive.
Choosing and using SSDs
Pick the drive for the work, check its health from Linux, and keep TRIM running. Hard disks still win on cost per terabyte for bulk data.
In simple words
Use an SSD wherever speed matters: the OS, apps, databases, build folders and laptops. Use hard disks for big, cold data such as backups and media archives, where price per terabyte counts more than speed. Check drive health now and then, and keep 10 to 20% free.
Think of it like
A sports car and a lorry. The sports car (SSD) gets you anywhere fast. The lorry (hard disk) is slow but moves enormous loads cheaply. A good fleet has both and uses each for what it is good at.
Swipe sideways to see the whole diagram
- 1A hard disk must seek the arm and wait for the platter to rotate before each random read.
- 2An SSD reaches any page electrically, so random access is almost as fast as sequential.
| Workload | Good choice | Why |
|---|---|---|
| Laptop, OS, apps | NVMe TLC with DRAM or HMB | Fast boot and launch, low power |
| Database, PostgreSQL | Enterprise NVMe with PLP | Low latency fsync, steady writes, power safety |
| Build server, CI | NVMe TLC, high TBW | Many small writes |
| Media library, read-heavy | QLC SSD or hard disk | Cheap capacity, few rewrites |
| Backups, archives | Hard disks or tape | Lowest cost per TB |
| Gaming | NVMe PCIe 4.0 | Fast asset streaming |
Check your drives from Linux
lsblk -d -o NAME,ROTA,TRAN,SIZE,MODEL # ROTA 0 = SSD, 1 = spinning disk
sudo apt install -y nvme-cli smartmontools
sudo nvme list
sudo nvme smart-log /dev/nvme0 # wear, temperature, errors
sudo smartctl -a /dev/nvme0n1
systemctl status fstrim.timer # weekly TRIM
sudo fstrim -av # TRIM now, verbose$ sudo nvme smart-log /dev/nvme0 critical_warning : 0 temperature : 41 °C available_spare : 100% percentage_used : 3% data_units_written : 41,982,310 (21.49 TB) media_errors : 0 unsafe_shutdowns : 12
| SMART field | Read it as |
|---|---|
percentage_used | Share of rated endurance consumed. Plan replacement as it nears 100%. |
available_spare | Spare blocks left. A falling value means blocks are being retired. |
data_units_written | Units of 512,000 bytes. Compare with the TBW rating. |
media_errors | Uncorrectable errors. Anything above 0: back up now. |
critical_warning | Non-zero means spare, temperature or read-only trouble. |
unsafe_shutdowns | Power cuts without a clean flush; matters for drives without PLP. |
Measure it yourself
fio --name=randread --filename=/tmp/fio.test --size=2G --direct=1 \
--rw=randread --bs=4k --iodepth=32 --numjobs=4 --runtime=30 --time_based --group_reportingTip: `--direct=1` bypasses the page cache, so you measure the drive and not RAM. Repeat with `--rw=randwrite` and `--iodepth=1` to see latency the way a database feels it.
Which one does what?
Every part of an SSD on one page: what it is, its everyday picture, and the term to search for.
| Part or idea | Its job | Think of it as | Technical name |
|---|---|---|---|
| Flash cell | Stores bits as trapped charge | A jar of marbles | Charge trap transistor |
| Bits per cell | Trades endurance for capacity | Telling water levels apart | SLC, MLC, TLC, QLC |
| Page | Smallest unit to write | A page of a book | 16 KiB NAND page |
| Block | Smallest unit to erase | A whole book | Erase block |
| 3D NAND | More layers, more capacity | Adding floors | Vertical NAND strings |
| Controller | Runs everything | Warehouse manager | SSD controller, firmware |
| FTL | Maps logical to physical | Hotel booking register | Flash translation layer |
| Garbage collection | Reclaims stale pages | Tearing out crossed out sheets | GC, TRIM, Deallocate |
| Wear leveling | Spreads erases evenly | Rotating tyres | Static and dynamic WL |
| ECC | Fixes drifting bits | A spell checker for bits | LDPC |
| SLC cache | Absorbs bursts | A walk-in fridge | pSLC write cache |
| DRAM / HMB | Holds the map | The manager's clipboard | DRAM cache, Host Memory Buffer |
| NVMe | Fast path to the CPU | A checkout per aisle | NVMe over PCIe |
| PLP | Survives power cuts | Backup generator | Power loss protection |
Next time a progress bar races and then slows down halfway through a big copy, you will know why: the SLC cache filled, the controller is folding data, and somewhere a few billion electrons are being tucked into place, one 16 KiB page at a time.