Two-Level Systems The physics of quantum hardware

IBM: The Transmon, and the Junction That Makes It Possible

Post 2 of Two-Level Systems, a 28-part series on the physics of quantum computing hardware. Start with the primer if you haven’t.

26 August 2026

Modality Superconducting circuit (transmon)
The qubit is The two lowest energy levels of a nonlinear microwave oscillator — physically, whether one extra Cooper pair’s worth of charge has sloshed across a Josephson junction
Lives at ~15 mK, in a dilution refrigerator, magnetically shielded
Largest public device Nighthawk, 120 qubits on a square lattice with 218 tunable couplers (Nov 2025). Heron r2, 156 qubits on a heavy-hex lattice, is the production workhorse
1Q / 2Q fidelity 2Q: reported EPLG (error per layered gate, a whole-device benchmark) of 9.93 × 10⁻⁴ on ibm_kingston and 3.98 × 10⁻³ on ibm_marrakesh — both Heron r2. Read that spread carefully; see below
Coherence Heron r2 medians around T₁ ≈ 180–195 μs, T₂ ≈ 95–120 μs
Gate time 1Q tens of ns; 2Q CZ on the order of 100 ns (the older cross-resonance gate ran 200–500 ns)
Connectivity Nearest-neighbour. Heavy-hex (degree 3) on Heron; square lattice (degree 4) on Nighthawk
Dominant error Two-level-system defects in oxides and interfaces, plus dephasing; at the device level, frequency crowding and crosstalk
Error correction Components demonstrated on Loon (112 qubits, Nov 2025). No below-threshold logical memory demonstrated. Targeting qLDPC codes, not the surface code
Source IBM Quantum published specifications and blog posts; Bravyi et al., Nature 627 (2024); device data from published experiments

A sandwich two nanometres thick

Every superconducting quantum computer in the world — IBM’s, Google’s, everyone’s — depends on one component. It is a sandwich: a strip of aluminium, an insulating layer of aluminium oxide about two nanometres thick, and another strip of aluminium on top. That’s it. That’s the Josephson junction.

Cool it below about one kelvin and the aluminium goes superconducting: electrons pair up into Cooper pairs and flow with exactly zero resistance. The oxide layer is an insulator, so classically nothing should cross it. Quantum mechanically, Cooper pairs tunnel through.

What makes this the single most important object in the field is a property that sounds technical and is actually everything: the Josephson junction is a nonlinear inductor that dissipates no energy.

Hold onto that, because the rest of this post follows from it.

Why you need the junction

Start with an ordinary LC circuit — an inductor and a capacitor. Energy sloshes between the magnetic field of the inductor and the electric field of the capacitor. It’s a harmonic oscillator, the same mathematics as a mass on a spring.

Cool it to millikelvin temperatures and quantum mechanics takes over. The oscillator’s energy is no longer continuous; it comes in discrete levels. You have a ladder of states — a microwave photon in the circuit, two photons, three — and you might reasonably think: call the ground state |0⟩, call the first excited state |1⟩, done.

It doesn’t work, and the reason is the defining problem of the whole platform.

A harmonic oscillator’s energy levels are evenly spaced. The gap from level 0 to level 1 is exactly the gap from 1 to 2, and from 2 to 3, forever. To do anything with the qubit you have to drive it with a microwave pulse tuned to the 0→1 transition. But that same pulse is perfectly tuned to 1→2, and to 2→3. You cannot excite the qubit from 0 to 1 and stop there. The population climbs straight up the ladder and out of your two-level subspace.

An evenly spaced ladder is not a qubit. It’s an oscillator, and you cannot isolate two levels of it no matter how clever your pulse is.

So you need to make the ladder uneven — anharmonic — so the 0→1 transition sits at a different frequency from 1→2. Then a pulse tuned to 0→1 is off-resonant with everything else, and you have addressed exactly two levels out of infinitely many.

To make an oscillator anharmonic, you need a nonlinear circuit element. And it must not dissipate energy, because dissipation is decoherence and decoherence destroys the qubit.

Energy level ladders: a harmonic LC oscillator has evenly spaced levels, so one pulse drives every transition; the transmon's junction makes the ladder uneven, detuning 1 to 2 from 0 to 1 by approximately the charging energy. Harmonic oscillator (LC circuit) Evenly spaced. Not a qubit. 0 1 2 3 same ω One pulse drives every step — population climbs the ladder. Transmon (junction + big capacitor) Anharmonic. Two levels addressable. 0 1 2 3 ω₀₁ detuned by ≈ −E_C A pulse at ω₀₁ is off-resonant with 1→2. Exactly two levels, isolated.
The whole problem, and the whole fix. Left: evenly spaced levels mean a pulse tuned to 0→1 also drives 1→2, and population climbs out of the qubit. Right: the Josephson junction’s nonlinearity detunes 1→2 by roughly −E_C, so a pulse at ω₀₁ addresses exactly two levels.

A resistor is nonlinear-ish and dissipates. A diode dissipates. Essentially every nonlinear component in ordinary electronics dissipates. The Josephson junction is nonlinear and dissipationless, and at low temperature it is effectively the only such element we have.

That is why superconducting qubits exist at all. Not because superconductivity is glamorous — because the Josephson junction is the one lossless nonlinearity nature offers.

The transmon, and the trade it makes

The junction gives you anharmonicity. It does not, by itself, give you a good qubit.

The first superconducting qubits — Cooper pair boxes, demonstrated at NEC by Nakamura and colleagues in 1999 and developed at Yale and Saclay — used a small junction and a small capacitance. In that regime the qubit’s frequency depends sharply on the number of Cooper pairs sitting on the island. That’s a feature for control and a catastrophe in practice: every stray charge fluctuation in the substrate — and there are many — shifts the qubit’s frequency. Coherence times were nanoseconds.

In 2007, Jens Koch and colleagues at Yale published the fix, in a paper whose title is the whole idea: Charge-insensitive qubit design derived from the Cooper pair box. Their proposal, the transmon, does something almost aggressively simple. Shunt the junction with a large capacitor.

The physics is a competition between two energies. The Josephson energy E_J measures how readily Cooper pairs tunnel; the charging energy E_C is the cost of putting one more pair on the island. Their ratio decides the qubit’s character.

Koch’s insight was that as you increase E_J/E_C, two things happen at very different rates:

  • Sensitivity to charge noise falls off exponentially, roughly as exp(−√(8E_J/E_C)).
  • Anharmonicity — the thing you needed the junction for — falls off only polynomially, ending up approximately equal to −E_C.

Exponential versus polynomial. Push E_J/E_C to about 50 or more and charge noise is suppressed by orders of magnitude while you retain a few hundred MHz of anharmonicity — not much, but enough to address two levels if your pulses are shaped carefully.

That asymmetry is the whole design. The transmon is a deliberately bad Cooper pair box: it gives up most of its nonlinearity in exchange for near-immunity to the noise that was killing its predecessor. Coherence went from nanoseconds to, eventually, hundreds of microseconds — a factor of about a hundred thousand.

The price is permanent, and it echoes through everything IBM does. A transmon’s anharmonicity is small, typically 200–300 MHz against a qubit frequency around 5 GHz — a few percent. Your control pulses must be gentle enough not to drive 1→2 leakage, which sets a floor on how fast you can gate. And crucially, transmons in a lattice all have similar frequencies, so they interfere with each other. Frequency crowding is the transmon’s original sin, and most of IBM’s architectural history is a response to it.

Deeper: the transmon Hamiltonian

A Josephson junction shunted by a capacitor is described by

H = 4E_C \hat{n}^2 - E_J \cos\hat{\varphi}

where n̂ counts Cooper pairs transferred across the junction and φ̂ is the superconducting phase difference. The first term is kinetic, the second is potential.

Everything hinges on that cosine. Expand it for small φ: cos φ ≈ 1 − φ²/2 + φ⁴/24. The quadratic term gives a harmonic oscillator with ω₀₁ ≈ √(8E_J E_C)/ħ. The quartic term is the anharmonicity — the entire reason the device is a qubit — and it yields a 1→2 transition detuned from 0→1 by approximately −E_C.

So E_C does double duty, and in opposite directions: large E_C gives you the anharmonicity you need, and small E_C (relative to E_J) gives you the charge insensitivity you need. The transmon sits where that tension is least bad. Every superconducting qubit in later posts — the unimon, the fluxonium, the cat qubit — is a different answer to the same tension.

Controlling one qubit

A transmon’s 0→1 transition sits around 5 GHz, in the microwave band. That is a convenient accident of engineering history: this is the frequency range in which mobile phones and satellite links operate, so the signal generators, amplifiers, mixers, and cabling already exist and are excellent.

To rotate a single qubit, send a microwave pulse down a control line at the qubit’s frequency. The qubit undergoes Rabi oscillation between |0⟩ and |1⟩; how far it rotates depends on the pulse amplitude and duration, and the axis of rotation depends on the pulse phase. A pulse of the right area is an X gate. Half that is a √X. Change the phase and you rotate about a different axis. Single-qubit gates take tens of nanoseconds and are routinely accurate to better than one part in ten thousand.

The complication is that small anharmonicity. A short pulse has a broad frequency spectrum — this is just Fourier — so a fast pulse tuned to 0→1 has spectral weight at 1→2 as well, and leaks population out of the computational subspace. Leakage is worse than an error: an error is something error correction can fix, while a leaked qubit is no longer a qubit and the code doesn’t know what to do with it.

The standard fix is a pulse-shaping technique called DRAG, which adds a component proportional to the derivative of the pulse envelope, tuned to cancel the leakage amplitude. It’s a nice example of a theme that recurs across every platform in this series: a physical limitation absorbed by control engineering rather than solved in hardware.

Readout works differently. Each qubit is coupled to a small microwave resonator, detuned far enough that no energy is exchanged. In that dispersive regime the resonator’s frequency shifts slightly depending on the qubit’s state. Bounce a probe tone off the resonator, measure the phase of what comes back, and you learn the qubit’s state without directly touching it. The returning signal is a few microwave photons and needs a near-quantum-limited amplifier — usually a Josephson parametric amplifier, the same junction physics used a third way.

Making two qubits talk — and IBM’s most interesting reversal

Single-qubit gates are the easy part. The two-qubit gate is where superconducting architectures actually differ, and IBM’s history here is the most instructive thing about the company.

For roughly a decade, IBM made a distinctive bet: all-fixed-frequency transmons.

The alternative, used by Google and others, is to make qubits frequency-tunable by putting a loop with two junctions (a SQUID) in the circuit, so magnetic flux tunes the qubit’s frequency. Tunability makes two-qubit gates easy — slide two qubits into resonance, let them interact, slide them apart. But it opens a new decoherence channel: the qubit’s frequency now depends on magnetic flux, and magnetic flux noise is everywhere. Tunable qubits generally dephase faster.

IBM refused the trade. Fixed-frequency transmons are quieter, but if you can’t tune anything, how do two qubits interact?

The answer is the cross-resonance gate: drive qubit A at qubit B’s frequency. Because they are weakly coupled, the drive reaches B through A, and the resulting interaction depends on A’s state — which is exactly the conditional dynamics an entangling gate requires. With echo sequences to cancel unwanted terms, it produces a CNOT. No flux, no tuning, no flux noise.

It worked, and IBM built its reputation on it. But it has two costs. Cross-resonance is slow — 200 to 500 nanoseconds, against tens of nanoseconds for a single-qubit gate — and it demands that neighbouring qubits sit in specific frequency relationships. Since the frequency of a fabricated transmon can’t be tuned after the fact, you must manufacture every qubit to its target frequency. Get a few wrong and you have frequency collisions: pairs whose gate simply doesn’t work well. On a large chip, yield becomes a frequency-allocation problem.

IBM’s response was architectural. The heavy-hex lattice — the distinctive layout on Heron and its predecessors — gives each qubit at most three neighbours instead of four. Fewer neighbours means fewer frequency constraints per qubit, fewer collisions, and less crosstalk. The cost is sparser connectivity: circuits need more SWAP operations to bring distant qubits together, and every SWAP is three CNOTs’ worth of error.

Then, with Heron in 2023, IBM changed its mind.

Heron keeps fixed-frequency qubits but inserts a tunable coupler between each pair: a separate circuit element, flux-tunable, sitting between two static qubits. Now the coupling turns on and off while the qubits themselves stay put. You get switchable interactions without making the qubits flux-sensitive, the native gate becomes a fast CZ instead of an echoed cross-resonance, and — the real prize — you can switch the coupling off, suppressing the always-on residual interaction that had been quietly limiting fidelity for years.

The consequences show up immediately in the roadmap. Heavy-hex existed to manage frequency collisions in a fixed-coupling world. With tunable couplers, that constraint relaxes — and Nighthawk, announced in November 2025, goes back to a square lattice: 120 qubits, degree four, 218 tunable couplers, and a stated capacity of 5,000 two-qubit gates in a circuit.

The lesson worth taking is not “IBM was wrong.” Fixed-frequency all-microwave control was a defensible bet that produced a decade of real results. The lesson is that the lattice geometry was never fundamental — it was a workaround for a specific gate mechanism, and it changed when the mechanism did. Watch for this in every post in this series: much of what looks like deep architecture is scar tissue from one component choice.

How it fails

Transmons are among the best-characterized qubits in existence, so the failure modes are known in unusual detail.

Energy relaxation (T₁). The excited state decays. The dominant culprit, at IBM’s level of materials quality, is two-level-system defects — microscopic imperfections in the aluminium oxide of the junction, in the surface oxides of the wiring, and at the interface with the substrate. Each defect is itself a little two-level quantum system, and when one happens to sit near the qubit’s frequency it absorbs the qubit’s energy.

There is an irony here worth naming, since this series borrows the term for its title: the thing destroying your two-level system is another two-level system. TLS defects drift and reconfigure over hours, which is why a superconducting chip’s performance is not stable day to day and why IBM recalibrates continuously.

Other T₁ channels: quasiparticles (broken Cooper pairs, generated by stray infrared light or cosmic rays), Purcell decay into the readout resonator, and coupling into unintended modes of the package.

Dephasing (T₂). On Heron r2, T₂ lands around 95–120 μs against T₁ of 180–195 μs. Recall from the primer that T₂ ≤ 2T₁; these devices sit well below that ceiling, which tells you dephasing is doing real independent damage rather than merely tracking energy loss.

The ratio that matters. T₁ ≈ 190 μs with 100 ns two-qubit gates gives roughly 2,000 gate times per coherence time. That’s the honest figure of merit — and it is nowhere near what an uncorrected useful algorithm needs, which is the entire reason error correction exists.

Read the fidelity numbers carefully. The stat block quotes EPLG of 9.93 × 10⁻⁴ on ibm_kingston and 3.98 × 10⁻³ on ibm_marrakesh. Both are Heron r2. Same generation, same design, same fab, roughly a fourfold difference in error rate.

That spread is not a scandal; it is what device-to-device variation in a solid-state platform actually looks like, and it is the concrete version of the warning from the primer. When a company quotes one number, ask which device, on which day, by which benchmark. EPLG is a deliberately honest metric — it measures errors per layer of gates executed in parallel across the whole chip, so it captures crosstalk that pairwise randomized benchmarking hides. It is consequently a larger, less flattering number than the best-pair figures often quoted, and it is the one worth comparing.

The bet that actually matters: qLDPC over the surface code

The primer argued that error correction, not qubit count, is the metric. IBM’s most consequential decision is which code to build toward, and here they have diverged from nearly everyone.

The surface code is the default for good reason: it needs only nearest-neighbour connectivity, which is all a 2D chip naturally provides, and its threshold is a comfortable ~1%. Its weakness is overhead. Encoding one logical qubit takes something like a thousand physical qubits at realistic error rates.

In 2024, Sergey Bravyi and colleagues at IBM published a Nature paper proposing bivariate bicycle codes, a family of quantum low-density parity-check codes. Their headline example, nicknamed the gross code (144 is a dozen dozen), is a [[144,12,12]] code: 12 logical qubits in 144 data qubits plus 144 check qubits — 288 physical qubits total.

Twelve logical qubits from 288 physical. Compare with a surface code of similar protective distance, which would need thousands for the same twelve. The paper puts the saving at better than ten to one, at a circuit-level threshold of roughly 0.7–0.8% — competitive with the surface code.

There is, inevitably, a catch, and it is a physics catch rather than an engineering inconvenience. The gross code’s stabilizer checks have weight six, and of the six connections each qubit needs, two are non-local — they reach across the chip rather than to a neighbour.

The surface code’s great virtue was that it asked nothing of your hardware beyond a grid. qLDPC codes ask for long-range wiring. On a planar superconducting chip, where every connection is a physical piece of metal that can also carry noise and host defects, that is a serious demand.

So IBM has made an explicit trade: buy a factor of ten in qubit overhead by paying in wiring complexity. That is what Loon is for. Announced alongside Nighthawk in November 2025, Loon is a 112-qubit experimental device whose purpose is not performance but demonstrating the components qLDPC requires — multi-layer routing on chip, long-range couplers connecting distant qubits, and fast qubit reset. It is a physics testbed for the wiring problem.

Whether that trade pays is genuinely open. It is the most interesting unresolved question in superconducting quantum computing, and it will take years of hardware to settle.

Demonstrated vs. proposed

Separating the two, since IBM publishes an unusually detailed roadmap and the distinction gets lost in coverage.

Demonstrated and published: - Transmons with T₁ around 180–195 μs at 156-qubit scale, in production, publicly accessible. - Tunable couplers with fast CZ gates, and EPLG under 10⁻³ on the best Heron r2 devices. - Nighthawk: 120 qubits, square lattice, 218 tunable couplers (Nov 2025). - Loon: the individual hardware components qLDPC needs, on a 112-qubit device (Nov 2025). - The gross code, as theory with numerical simulation, peer-reviewed in Nature.

Proposed, on the roadmap, not yet demonstrated: - Quantum advantage by end of 2026 — a claim that depends heavily on what counts as advantage, and one where previous claims across the industry have repeatedly been met with classical simulations that caught up. - Starling, 2029: 200 logical qubits, 100 million gates. - Blue Jay, 2033: 2,000 logical qubits. - A below-threshold logical memory on IBM hardware. This is the one to watch. IBM has not yet published the experiment Google published in 2024 — logical error rate falling as code distance grows. Until a qLDPC code is run on real hardware and shown to improve with size, the ten-to-one overhead advantage remains a simulation result. A well-founded one, but a simulation.

That last gap is the honest summary of where IBM stands: the best-documented transmon hardware in the world, a genuinely promising and peer-reviewed code, and the connection between them still to be built.

The receipts

  • J. Koch et al., “Charge-insensitive qubit design derived from the Cooper pair box,” Phys. Rev. A 76, 042319 (2007), arXiv:cond-mat/0703002. The transmon. Yale, from the Schoelkopf and Girvin groups. Nearly every qubit in this arc is a descendant.
  • Y. Nakamura, Yu. A. Pashkin, J. S. Tsai, “Coherent control of macroscopic quantum states in a single-Cooper-pair box,” Nature 398, 786 (1999). The ancestor — the first demonstration that a fabricated circuit could behave as a coherent quantum two-level system.
  • C. Rigetti and M. Devoret, “Fully microwave-tunable universal gates in superconducting qubits with linear couplings and fixed transition frequencies,” Phys. Rev. B 81, 134507 (2010). The cross-resonance gate, from the person who later founded Rigetti — post 4.
  • J. M. Chow et al., “Simple all-microwave entangling gate for fixed-frequency superconducting qubits,” Phys. Rev. Lett. 107, 080502 (2011). Cross-resonance in practice at IBM.
  • S. Bravyi, A. Cross, J. Gambetta, D. Maslov, P. Rall, T. Yoder, “High-threshold and low-overhead fault-tolerant quantum memory,” Nature 627, 778 (2024). The gross code. The most consequential paper IBM has published in years.
  • P. Krantz et al., “A Quantum Engineer’s Guide to Superconducting Qubits,” Applied Physics Reviews 6, 021318 (2019), arXiv:1904.06560. Not IBM, but the best single document on how these circuits actually work. If one thing here interested you, read this next.

Next

Post 3: Google Quantum AI. Same physical qubit, opposite choice at nearly every decision point — tunable qubits instead of fixed, a square lattice instead of heavy-hex, and the surface code instead of qLDPC. Google also did the experiment IBM hasn’t: error correction demonstrably below threshold, with the logical error rate falling as the code grew.


This series explains the physics of quantum computing hardware. It is not investment advice, and it does not rank or evaluate companies as businesses.