In brief

Modern AI clusters connect tens of thousands of GPUs: between them flow colossal data streams at throughputs that copper can no longer sustain without burning tens of megawatts in pure overhead. Co-Packaged Optics (CPO) is an architectural break: instead of placing optical transceivers in separate pluggable modules, they are physically integrated into the same package as the switching ASIC or accelerator. The result: electrical losses drop from 22 dB to 4 dB, per-port power falls by more than 70%, and the bandwidth density available at the chip edge increases by an order of magnitude. Broadcom, Nvidia, Ayar Labs, Lightmatter, and Celestial AI have all announced CPO products in 2025-2026. The market is estimated at $20 billion by 2036, with a 37% annual growth rate.


The copper wall: a physical problem, not a software one

Imagine a hospital whose corridors were sized for 1990s traffic. Doctors, nurses, and gurneys cross paths, block each other, wait. Expanding the operating rooms changes nothing if the corridors stay narrow. That is exactly the situation of AI datacenters in 2025: GPUs are becoming more powerful, but the infrastructure connecting them — copper interconnects — creates bottlenecks that raw compute cannot compensate.

Three physical limits converge:

Signal attenuation at high frequencies. Beyond 112 gigabits per second per lane, electrical losses in printed circuit boards and connectors explode. On a 200 Gb/s/lane link — the speed of current CPO products — loss in a classic pluggable transceiver system reaches 22 dB. An integrated CPO system reduces that loss to 4 dB. Source: NVIDIA Technical Blog, 2025.

Power consumption of pluggable transceivers. A standard 800G OSFP transceiver consumes 15 to 30 watts. For a cluster of one million GPUs (each requiring six network ports), the power dedicated to transceivers alone exceeds 180 megawatts — an entire power plant for optical plumbing. Source: SemiAnalysis 2025, tokenring/FinancialContent.

Bandwidth density at the chip edge. Copper connections are physically limited by the chip perimeter (the “shoreline”). CPO solutions achieve 1 to 2 terabits per second per millimeter of shoreline in 2025. Source: EE Times OFC 2025.

At the scale of today’s AI racks approaching 100 kilowatts of density, interconnects consumed up to 30% of total cluster power in early 2025. CPO aims to reduce this “optical tax” by more than 80% according to late-2025 pilot deployments. Source: tokenring/wral.

In plain terms: copper works well up to a certain speed. Beyond that, every additional gigabit per second costs more and more energy to compensate for losses. CPO sidesteps this wall by replacing electrons with photons — particles of light that do not dissipate in copper.


Silicon photonics: manufacturing optics like silicon

Before understanding CPO, one must grasp the foundational building block that makes it possible: silicon photonics (SiPh).

The idea is compelling in its simplicity: use the same manufacturing processes as for conventional electronic chips (CMOS technology) to engrave optical components directly in silicon. Instead of transistors switching electrons, one fabricates waveguides routing photons, modulators encoding data on a light beam, and photodetectors converting light back to an electrical signal.

Vocabulary

  • SiPh (Silicon Photonics): technology for manufacturing optical components — waveguides, modulators, photodetectors — integrated into a silicon chip via standard CMOS processes.
  • PIC (Photonic Integrated Circuit): an integrated photonic chip, the optical equivalent of an electronic integrated circuit.
  • Lane: an individual transmission channel — one track of a multi-channel link. “200 Gb/s per lane” means 200 gigabits per second on each track.
  • Shoreline: the perimeter of a chip available for incoming/outgoing connections. Bandwidth density is measured in Tbps/mm of shoreline.

Key photonic components

Four building blocks constitute a SiPh system for a datacenter:

Modulators encode digital data onto a continuous light beam. Two families compete:

  • Mach-Zehnder modulators (MZM): thermally robust, wide bandwidth, but more energy-hungry (5 to 10 pJ/bit). Used by early Broadcom and Intel generations.
  • Micro-ring modulators (MRM): highly efficient (1 to 2 pJ/bit for the modulator alone), but sensitive to temperature variations — a few degrees of shift detunes the ring resonance and cuts the signal. Nvidia Quantum-X adopts them for their energy efficiency, at the cost of precise active thermal control.

Lasers pose a particular challenge: silicon is a poor light emitter (indirect electronic bandgap). Two strategies coexist:

  • External laser: a separate light source, injected into the photonic chip via a coupler. Advantage: it can be replaced without dismantling the package. Example: Ayar Labs’ SuperNova, which emits 16 simultaneous wavelengths via WDM multiplexing.
  • III-V heterointegration: direct bonding of InP or GaAs materials — which emit light readily — onto the silicon wafer. Active research track, not yet in mass production.

Photodetectors convert light into electrical current. Germanium-on-silicon (Ge-on-Si) is the dominant solution: CMOS-compatible, efficient in the 1310-1550 nm spectral window used by CPO systems.

Waveguides, couplers, and multiplexers form the optical plumbing: they route light from one component to another, split beams, and separate wavelengths.

TSMC COUPE: the foundry enters the game

TSMC has developed the COUPE (Compact Universal Photonic Engine) platform, a PIC integration system on CoWoS (Chip on Wafer on Substrate) substrate. The architecture consists of three components: COUPE (logic ASIC integrated on photonics via SoIC), COI (Complementary Optical Interconnect), and iFAU (Integrated Fiber Array Unit). COUPE qualification for advanced pluggable systems occurred in 2025; full integration in CPO mode on CoWoS is planned for 2026. Sources: TrendForce/SEMICON Taiwan, September 2025; 3D InCites, October 2025.

In plain terms: silicon photonics fabricates “printed circuits for light” by reusing chip manufacturing plants. That is what makes CPO economically viable: no need for an entirely new manufacturing chain.


CPO: integrating optics into the package

Silicon photonics is the technology. Co-Packaged Optics (CPO) is the packaging architecture that exploits it. Understanding the difference between the three generations of optical architecture is essential for reading vendor announcements.

Vocabulary

  • Pluggable optics: optical transceiver in a standardized module (QSFP, OSFP) that plugs into a front-panel port of the switch. Hot-swappable, but connected to the ASIC via a long electrical path on the PCB.
  • NPO (Near-Package Optics): optical engine mounted on the board a few centimeters from the ASIC, outside the package. Reduces PCB losses but does not eliminate the electrical penalty of crossing the package boundary.
  • CPO (Co-Packaged Optics): optical engine physically integrated into the ASIC package, on the same substrate — silicon or organic interposer. Ultra-short electrical connection between the ASIC and the optical engine.
GenerationArchitectureElectrical lossPower / 800G port
Pluggable optics (OSFP)External transceiver~22 dB15–30 W
Near-Package OpticsEngine on board~10 dB8–12 W
CPOEngine inside the package~4 dB3.5–9 W

Sources: APNIC Blog (CPO deep dive, May 2025); NVIDIA Technical Blog (2025); Broadcom press release (TH6, October 2025).

pJ/bit metrics: reading vendor figures

The picojoule per bit (pJ/bit) is the standard energy efficiency unit for interconnects. A system consuming 1 pJ/bit uses 1 picojoule of energy to transmit one bit of information. Lower is better.

Triangulated values from multiple sources:

TechnologypJ/bitMeasurement scopeSource
Pluggable optics 800G (OSFP)~15–20 pJ/bitFull portSemiAnalysis, APNIC Blog
Broadcom TH5-Bailly CPO (51.2 Tbps)~5.5 pJ/bitFull portAPNIC Blog
Broadcom TH6-Davisson CPO (102.4 Tbps)<3.8 pJ/bitFull portBroadcom press release, StorageReview
NVIDIA Quantum-X (MRM modulators only)~1–2 pJ/bitModulator onlyAPNIC Blog
Long-term industry target<1 pJ/bitFull portIDTechEx, 2028+ horizon [roadmap]

An important nuance: pJ/bit values announced by vendors sometimes refer only to the optical modulator, excluding the laser, electronic driver, and DSP. The energy cost of a full port is systematically higher. A fair comparison covers the full port, not the isolated component.

In plain terms: CPO does not simply replace copper with light. It eliminates the long electrical paths on printed circuit boards — the primary source of losses — by placing the optical transceiver a few millimeters from the ASIC, inside the same package.


The players: who is building CPO in 2025-2026

Five players have announced concrete CPO products, with radically different architectures. This is not a race to who arrives first, but to who delivers the best integration for each use case.

Broadcom — Tomahawk 6 Davisson: CPO for Ethernet switches

Broadcom is the industrial pioneer of CPO on Ethernet switches, with two generations already deployed.

TH5-Bailly CPO (2024, 51.2 Tbps): second generation, serves as the baseline reference for sector comparisons at ~5.5 pJ/bit.

TH6-Davisson CPO (announced October 2025, shipping): third generation. 102.4 terabits per second total bandwidth, 200 Gbps per lane. Power per 800G port: ~3.5 W, versus ~5.4 W for TH5 CPO and more than 9 W for a classic pluggable — a reduction exceeding 70%. Key innovation of TH6: laser modules are field-replaceable from the front panel, without dismantling the switch. This is a major operational break from previous generations where a laser failure required replacing the entire equipment. Sources: Broadcom press release, NextPlatform, StorageReview.

A fourth Broadcom generation at 400 Gbps per lane is in development. [roadmap, vendor claim]

Nvidia — Quantum-X Photonics and Spectrum-X Photonics

Nvidia enters CPO through its own InfiniBand and Ethernet switches, announced at GTC 2025.

Quantum-X Photonics (InfiniBand scale-out, available early 2026): 115 Tbps total bandwidth, 144 ports × 800 Gb/s. Architecture: Quantum-X800 ASIC (TSMC 4N process, 107 billion transistors) paired with six optical components integrating 18 photonic engines. Uses micro-ring modulators at 200 Gbps per lane. Efficiency: 3.5 times less power than a switch with pluggable transceivers; 9 W per CPO port versus 30 W per pluggable port. Source: NVIDIA Newsroom, NVIDIA Technical Blog.

Spectrum-X Photonics (Ethernet scale-out, H2 2026 [roadmap]): two models — SN6810 at 102.4 Tbps (128 × 800G ports) and SN6800 at 409.6 Tbps (512 × 800G ports). Source: NVIDIA Newsroom.

Nvidia also announced at GTC 2026 CPO for NVLink scale-up with the Feynman NVLink 8 CPO switches in 2028. Jensen Huang noted: “There will also be copper in scale up.” Both technologies will coexist. Source: HPCwire, March 2026. [roadmap, vendor claim]

Ayar Labs — TeraPHY: the UCIe optical chiplet

Founded in 2016 as an MIT spin-off, Ayar Labs takes a different strategy: not building a CPO switch, but an optical chiplet interoperable via the UCIe (Universal Chiplet Interconnect Express) interface that any chip designer can integrate into their design.

TeraPHY: 8 Tbps optical chiplet using micro-ring resonator modulators. Light source: SuperNova, an external laser emitting 16 WDM wavelengths simultaneously, separable from the package to facilitate maintenance.

At OFC 2025, Ayar Labs announced the “world’s first UCIe optical chiplet” reaching 8 Tbps. A notable partnership with Alchip (demonstration at TSMC OIP 2025): more than 100 Tbps of scale-up bandwidth per accelerator, more than 256 optical scale-up ports per device — targeting clusters of 1,000 to 10,000 GPUs in a single optical domain. Sources: Ayar Labs press release, Semiwiki, EE Times OFC 2025.

In March 2026, Ayar Labs raised $500 million to transition to mass production of CPO chiplets, with investors including AMD, MediaTek, and Alchip. A partnership with Wiwynn was also announced for rack-scale optically connected systems. Sources: The Register March 2026, Semiconductor Today March 2026.

Lightmatter — Passage: the active photonic interposer

Founded at MIT in 2017, Lightmatter adopts the most ambitious architecture: not a CPO transceiver, but an active 3D photonic interposer that sits between logic chiplets as an integrated optical interconnect fabric.

Passage M1000 (announced March 2025, available summer 2025 [early]): “3D Photonic Superchip”, active multi-reticle interposer enabling the assembly of large dies. Total optical bandwidth: 114 Tbps. Presented at Hot Chips 2025. Sources: Lightmatter press release, ServeTheHome Hot Chips 2025.

Passage L200 and L200X (2026): 3D CPO solutions targeting hyperscaler XPUs and switches. Announced bandwidths: 32 Tbps (L200) and 64 Tbps (L200X). Lightmatter claims the equivalent of 40 pluggable transceivers in a single L200 package, and a 2.7× reduction in training time versus copper. Source: Lightmatter press release. [vendor claim — not independently triangulated]

Celestial AI / Marvell — Photonic Fabric: optics inside the die

Celestial AI, founded around 2020, pushes integration even further: optics are integrated inside the die itself (in-die), not on a separate interposer.

Photonic Fabric Module (presented Hot Chips August 2025 [early]): first SoC with in-die optical interconnect. Architecture: TSMC 5nm ASIC (8 Tbps switch + HBM3e controllers + DDR5) + photonic PIC interposer + 2 HBM3e stacks + FAU — 2.5D package. Optical connectivity: 7.2 Tbps. A scale-up chiplet version announces 16 Tbps in a single chiplet, or “10× the capacity of current 1.6T scale-out ports” [vendor claim]. Sources: ServeTheHome Hot Chips 2025, Chiplet Marketplace.

Acquisition by Marvell (late 2025): Marvell acquires Celestial AI for ~$3.25 billion (cash + stock). Meaningful revenue expected in the second half of fiscal 2028, with a target run rate of $1 billion in Q4 FY2029. Integration into Marvell’s custom silicon portfolio for hyperscalers provides a direct path to deployment. Sources: Marvell investor relations, Optics.org.

In plain terms: five players, five architectures. Broadcom integrates CPO into its Ethernet switches. Nvidia integrates it into its InfiniBand and Ethernet switches. Ayar Labs offers an optical chiplet that other chips can embed. Lightmatter builds a 3D optical interposer. Celestial AI (now Marvell) places optics directly inside the die. These approaches are not strictly competing — they target different levels of the interconnect hierarchy.


Two distinct use cases: scale-out and scale-up

Understanding real CPO deployments requires distinguishing two levels of interconnection in AI datacenters, which have different needs.

Scale-out: connecting racks together

Scale-out refers to connections between nodes — between servers in a rack, between racks in a hall, between halls in a datacenter. This is the territory of network switches (Ethernet, InfiniBand). Typical distances range from a few meters to a few hundred meters.

This is the most mature market for CPO in 2025-2026. Broadcom TH6-Davisson and Nvidia Quantum-X/Spectrum-X Photonics switches directly target this segment. On a 10,000-GPU cluster, dropping from 30 W/port to 9 W/port on network switches represents a saving of approximately 25 to 30 megawatts — an order of magnitude that transforms the economic and environmental viability of mega-clusters.

Scale-up: connecting GPUs within the same domain

Scale-up refers to connections within a single node or between a small number of tightly coupled nodes — typically NVLink between GPUs in the same server or fabric. Distances are much shorter (from a few centimeters to a few meters), but required throughputs are extreme.

This is the territory of Ayar Labs and Lightmatter. The Alchip + Ayar Labs partnership targets accelerators directly integrating the TeraPHY chiplet, enabling bandwidth exceeding 100 Tbps per accelerator over distances of up to 100 m. Lightmatter Passage M1000 is presented as enabling “connecting thousands of GPUs in a single optical domain.”

In 2028, Nvidia plans to introduce CPO into its NVLink scale-up switches (Feynman NVLink 8 CPO), while maintaining some copper connections — acknowledging that optics does not replace copper for all distances. [roadmap, vendor claim] Source: HPCwire March 2026.

CriterionNVLink 4.0 (H100, electrical)CPO scale-up (Ayar/Lightmatter, 2026-2028)
BW per node900 GB/s (8 GPUs)>100 Tbps per accelerator [vendor claim]
Distance~1 m (copper cables)cm to hundreds of meters
MaturityMass productionPOC / early deployment [early]

In plain terms: scale-out CPO (in network switches) is commercial in 2025-2026. Scale-up CPO (in GPU/accelerator packages) is in demonstration and early deployment phase, with industrial maturity expected from 2027-2028.


Technical challenges and remaining hurdles

CPO is not a mature technology in the industrial sense. Several technical challenges condition its mass deployment.

Thermal management of optical components

Micro-ring resonator modulators and lasers are temperature-sensitive. When the neighboring ASIC heats under load, the optical components must be maintained in a narrow temperature range (a few degrees) to keep ring resonance stable. This requires active thermal control systems (PID loops, Peltier actuators) that add complexity and consume additional power.

Fiber coupling at scale

Aligning hundreds or thousands of optical fibers on a photonic interposer with tolerances of ±1 micrometer is a volume manufacturing challenge. TSMC’s iFAU (Integrated Fiber Array Unit) aims to standardize this coupling, but production yield at volume remains a critical non-public data point. Source: TSMC COUPE architecture, TrendForce 2025.

Laser replacement in case of failure

The laser is the optical component with the most limited lifespan. In a classic CPO architecture, a laser failure would require replacing the entire ASIC — prohibitive. Broadcom TH6-Davisson partially solves this problem by making laser modules field-replaceable from the front panel. Ayar Labs takes a different approach: the SuperNova external laser remains separate from the TeraPHY chiplet, facilitating replacement without touching the ASIC.

Manufacturing yield

Integrating photonic components (sensitive to surface defects and waveguide roughness) and advanced CMOS logic on the same substrate increases yield complexity. TSMC’s CoWoS + COUPE architecture separates photonic and logic processes to maximize overall yield — photonic chips and logic chips are manufactured separately, then assembled. Source: 3D InCites, October 2025.

CoWoS capacity at TSMC

TSMC’s advanced CoWoS packaging capacity was under strong tension in 2024-2025 (H100/H200/Blackwell demand). The expansion planned for 2026 should free up capacity for CoWoS CPO configurations, but lead times remain a real deployment constraint. Source: tokenring.

In plain terms: CPO is not “just” replacing a cable with fiber. It is a complex integration chain — photonic manufacturing, precise thermal management, precise assembly, volume yield — that the industry is in the process of mastering, but which will take another 2 to 3 years to be fully industrialized.


Deployment timeline: where things really stand

ProductPlayerStatusPeriod
TH5-Bailly CPO (51.2 Tbps)BroadcomDeployed (gen 2)2024
TH6-Davisson CPO (102.4 Tbps)BroadcomShippingOct 2025
COUPE pluggable qualificationTSMCQualified2025
COUPE CPO CoWoS integrationTSMCQualification ongoing2026
Quantum-X PhotonicsNvidiaCommercialEarly 2026
Ayar TeraPHY UCIe chipletAyar LabsFirst demonstrator [early]OFC 2025
Passage M1000 (114 Tbps)LightmatterAvailable [early]Summer 2025
Photonic Fabric ModuleCelestial AI / MarvellDemonstrator [early]Hot Chips Aug 2025
Spectrum-X PhotonicsNvidiaIn productionH2 2026 [roadmap]
Passage L200 / L200XLightmatterAnnounced2026 [early]
Feynman NVLink 8 CPONvidiaAnnounced2028 [roadmap, vendor claim]
Broadcom gen 4 (400G/lane)BroadcomIn developmentTBD [roadmap, vendor claim]

The global CPO market is estimated at a 37% CAGR to reach $20 billion by 2036. Source: IDTechEx, Co-Packaged Optics 2026-2036 report.

In 2026, the general status according to EDN is straightforward: scale-out CPO (switches) is in qualification and early commercial deployment — “pretty natural for a leading-edge technology.” Scale-up CPO (in GPU/accelerator packages) remains in POC and early deployment phase for the majority of players.


What CPO does not replace

An important boundary: CPO for AI datacenters has nothing in common with long-distance telecom fiber optics. These two worlds use light but in radically different contexts:

  • Long-distance telecom fiber: single-mode DWDM (Dense Wavelength Division Multiplexing) links, Erbium amplifiers (EDFA), distances of 100 to 10,000 km, throughput per fiber of a few Tbps, optimized for cost per bit over very long distances.
  • Datacenter CPO: links of a few centimeters to a few hundred meters, integration in the chip package, optimized for bandwidth density and energy efficiency at rack or datacenter scale.

Similarly, CPO is distinct from the high-performance networking article, which covers NVLink (electrical scale-up interconnect within the node) and classic InfiniBand/Ethernet. CPO is the photonic layer that replaces electronic transceivers in those same links — it is a transverse technology, not an alternative to the network protocols themselves.

In plain terms: CPO is not a new network protocol. It is a physical transmission technology — a replacement of the physical layer — that can be applied to the same NVLink, InfiniBand, or Ethernet infrastructures, but using light instead of electrons on the most critical paths.


Which CPO architecture for which need?

ContextRecommendationWhy
Datacenter Ethernet switch, 800G transceiver replacementBroadcom TH6-Davisson CPOCommercial product deliverable in 2025-2026, field-replaceable laser, >70% energy saving.
InfiniBand scale-out switch for Nvidia clustersNvidia Quantum-X PhotonicsNative Nvidia ecosystem, available early 2026, 3.5× less power than pluggable.
Custom accelerator with interoperable CPO chipletAyar Labs TeraPHY (UCIe)Standard UCIe interface, replaceable external laser, 8 Tbps per chiplet, $500M funding March 2026.
Hyperscaler XPU with integrated 3D opticsLightmatter Passage or Marvell/Celestial AIMost ambitious integration architectures, best effort for scale-up, 2026-2027 maturity [early].
CPO purchasing decision in 2026Start with scale-out (switches)Scale-out CPO is commercial; scale-up CPO in GPU packages remains in early deployment phase.

A few heuristics for teams evaluating CPO:

Do not confuse modulator pJ/bit with full system pJ/bit. Press release figures often cite the modulator alone. For a fair comparison, require the full port figure (laser + driver + DSP + modulator).

TCO maturity is not yet established. SemiAnalysis notes that a CPO switch cannot replace just its defective transceivers — the entire switch must be replaced. Total cost of ownership depends on CPO laser failure rates over the switch lifetime, data that does not yet exist at volume.

The replaceable laser is a key variable. Broadcom TH6 and Ayar Labs TeraPHY have solved it differently; Nvidia and Lightmatter have less clear architectures on this point. For an operator, it is a major operational criterion.


Key takeaways for 2026

CPO is no longer a laboratory curiosity. Broadcom is shipping CPO switches at 102.4 Tbps. Nvidia is selling CPO InfiniBand switches commercially. Photonic startups have raised hundreds of millions of dollars to move to mass production. TSMC is industrializing the COUPE platform to integrate optics into its CoWoS packages.

The transition to all-optical in AI datacenters follows an inevitable logic: as required throughputs increase and clusters reach hundreds of thousands of GPUs, the power consumed by electrical interconnects becomes unsustainable. Reducing that expense by more than 70% per port is an economic and environmental advantage that imposes itself.

The 2028 horizon will see CPO penetrate scale-up — interconnects directly in GPU/XPU packages, where throughputs are most demanding and energy gains most significant. Until then, silicon photonics continues to decline in cost through volume, and the challenges of yield, thermal management, and laser replacement are being resolved industrially.