<!-- Slide number: 1 -->
OPTICAL INTERCONNECTION NETWORKS ON THE WAY FROM CONCEPT TO TECHNOLOGY
Instructor: Davide Bertozzi
Notes:
Good Afternoon everyone,
In this presentation, we are going to talk about Enginnering a Bandwidht-Scalable Optical Layer for
a 3D Multi-core Processor With Awareness of Layout Constraints.
<!-- Slide number: 2 -->
Key idea
Optical on-chip and chip-to-chip communication
Main driver:
High bandwidth density
Target metrics:
1 Tbps/link
<10mW/Gb/s/link
Price: <<0.1$ per Gbps

Main drivers:
High bandwidth communication
and/or reduce power/bit
Target metrics:
1 Tbps/link
<1mW/Gb/s/link
Price: <0.01$ per Gbps

Seamless scaling to off-chip communication
<!-- Slide number: 3 -->
Background
(typically) off-chip
Laser source
Optical signal carried to the chip via optical fiber




Tapered input


Silicon waveguide
Silicon waveguide
Optical OOK modulation
<!-- Slide number: 4 -->
Wavelength Division Multiplexing

Fibre or waveguide

Multi-wavelength optical input delivers inherent parallelism opportunities for on-chip communication
Potentially, a bit parallel electronic signal might be converted into a wavelength parallel optical signal
Array of modulators needed
<!-- Slide number: 5 -->
Optical Modulator

Typically based on a silicon microring resonator
Light can be coupled into the microring (hence preventing light propagation) or can be made insensitive to it (hence enabling light propagation)
Carrier injection into the microring resonator is used to put it on- vs. off-resonance with respect to the input wavelength

Power of 0.01 pJ/bit in best-in-class devices
+
0.1 pJ/bit in the electronic driver
+
0.15 pJ/bit for themal stabilization
Modulation rates of 10 Gbps are today feasible, and 12.5, 25 and 40 Gbps will be certainly within reach
<!-- Slide number: 6 -->
Silicon waveguide

Excellent light confinement
1.3 Tbps through a single waveguide proven
(B.Lee et al, IEEE PTL, 20, 398 (2008))
Propagation-related optical power loss less than 1.5 dB/cm
Bending loss < 0.001 dB/turn (bending Rardius = 3um)
The real issue when it comes to optical power loss is waveguide crossings!
<!-- Slide number: 7 -->
Waveguide crossing
With direct waveguide crossings, lateral confinement in the photonic waveguide is lost near the crossing, causing diffraction of the light
A sizable fraction of the light is radiated away (from -1.1 to -1.7 dB)
This generates non-negligible crosstalk (e.g., -9dB) in the arms of the crossing

Notes:
12% crosstalk
Up to 77% diffraction
<!-- Slide number: 8 -->
Waveguide crossing
There exist several waveguide intersection optimization techniques, that can result into crossing losses as low as 0.5 and 0.2 dB

MMI taper

Elliptical taper
Tapers have different area footprints. A trade-off between taper size and insertion loss needs to be considered
E.g., only 0.18 dB for a 10um crossing length
0.7 dB are today industry standard for crossing losses
Notes:
5 to 10%
Will be 1%
Std 15%
<!-- Slide number: 9 -->
Receiver
Silicon is not a light absorbing material
No photodetection in 1.2μm to 1.6μm
Ge photodiodes promising


A receiver sensitivity of -16 dBm has been reported at 12.5 Gbps
-20dBm can be expected from technology evolution
Notes:
0.025 mW
0.01 mW
<!-- Slide number: 10 -->
Optics closer to the processor?
Optics has made a long way from long-haul telecommunication networks to data centers and multi-chip systems

On-chip wires are by definition inexpensive..

..but don’t forget the hell of nanoscale physics!
Will it be able to penetrate deeper into smaller scale systems?
<!-- Slide number: 11 -->
#
Silicon chip

Multi-chip
Datacenter
<!-- Slide number: 12 -->
Datacenters

Exponential increase in the internet traffic
(streaming video, social networking, cloud computing)

<!-- Slide number: 13 -->
Displacing the copper cable
Key issues:
High power overhead from E/O and O/E conversions
Limited scalability
latency overhead

Internet
Data Center


Core
Aggregation



ToR
switch









Servers
As lane speeds move to 10 Gbps and beyond,
fiber-optic technologies are displacing copper-based solutions
Photonic (all-optical) switching holds promise for
Keeping up with bandwidth (density), latency and power requirements
dynamic reconfiguration based on instantaneous demands, cyclical patterns, or prediction
<!-- Slide number: 14 -->
Hybrid Architectures
Incremental upgrade of operating data centers with commodity switches, reducing the cost of the upgrade

cThrough (RiceU, CMU, Intel)
Electrical network
Optical network
Circuit switching dynamically configured for pairs of racks with high bandwidth demands
Suitable for bulky traffic that lasts long enough to compensate for the reconfiguration overhead
Scalability limited by the number of optical ports of the switch (e.g., 64)
<!-- Slide number: 15 -->
All-Optical Circuit Switched Networks
Publicité
Often rely on optical MEMS switches
Reconfiguration time can be up to 20-30 ms!
TOO BAD!


Targeting long-term bulky data transfers (e.g.,enterprise networks)
Optical circuit switching is (largely) data rate agnostic and extremely energy efficient (scaling without replacements)
no packet processing, ultra-low latency and power

<!-- Slide number: 16 -->
All-Optical Packet Switched Networks
Native optical packet switching has long been a goal of the optics community. However, a number of fundamental challenges leave this vision a breakthrough away from widespread commercial adoption.

Better matches networks where
The duration of a flow between two nodes is small
all-to-all connectivity is required



The lack of robust optical buffers and registers causes a design paradigm which is radically different from packet-switched electronic networks
- O/E and E/O conversions, deflection routing, packet loss, injection control, timing constraints, ..
<!-- Slide number: 17 -->
Current Challenges
The cost of optical switch components is currently a barrier to entry into the data center
Number of supported duplex ports in optical circuit switching should scale from hundreds to thousands or tens of thousands
Switching times should improve from tens of msecs (driven by requirements of telecom industry) to few hundreds of usecs (ideally below 100), possibly migrating from MEMS to SOA (silicon optical amplifier)-based solutions (for packet switching: order of nsecs)
Lower the insertion loss from 5 to below 2dB
Continued evolution of WDM optical transceivers (low power, large-distance span, nx25G bandwidth,..)
<!-- Slide number: 18 -->
The chip I/O Bottleneck
Higher on-chip bandwidths more off-chip communication
Off-chip bandwidth scales through pin count & signaling rate
Pin counts limited by packaging constraints, chip size, and crosstalk
Power scales badly with signaling rates

Memory InterfaceController
25.6 GB/s @ 3.2GHz
I/O Controller
25 GB/s @ 3.2GHz(inbound)
[Kistler et al., IEEE Micro 26 (3) 10–23 (2006)]
Notes:
<!-- Slide number: 19 -->
Off-chip Communications
Element Interconnect Bus(on-chip communications)
delivers nearly an order of magnitude more bandwidth:
205 GB/s @ 3.2 GHz

Memory InterfaceController
25.6 GB/s @ 3.2GHz
I/O Controller
25 GB/s @ 3.2GHz(inbound)
[Kistler et al., IEEE Micro 26 (3) 10–23 (2006)]
Notes:
<!-- Slide number: 20 -->
Can optical links help?
Through WDM,
Concurrently
transmit multiple
spectrally
parallel streams
of data through
a single optical
waveguide,


waveguide
Overcome routing congestion issues on a chip
Overcoming the I/O pin count limitation through increased bandwidth density
Overcoming the concern of bond pad capacity and pin inductance
CMOS-compatible solution for integrating high bandwidth-density off-chip optical I/O which can overcome some of these packaging limitations while adhering to
pJ/bit-scale power efficiency requirements will be within reach in a few years
<!-- Slide number: 21 -->
I/O Bandwidth Density
Microring based point-to-point optical link
Based on reported best of class devices
Separate optical dies with limited electronics within the package

Inter-channel crosstalk
suppression
Dense WDM
fiber
Front-end circuits
within the module
(flip chip bonding, or monolithic integration)
12.5 or 25 Gbps
Optimistic RX sensitivities:
-20 dBm at 12.5 Gbps
-16 dBm at 25 Gbps
Thermal tuning of microrings considered
Power budget: 20 dBm
Noam Ophir, Christopher Mineo, David Mountain, Keren Bergman, "Silicon Photonic Microring Links for High-Bandwidth-Density, Low-Power Chip I/O," IEEE MICRO, Jan.-Feb. 2013 (vol. 33 no. 1) pp. 54-67
Notes:
100 mW
<!-- Slide number: 22 -->
THE CONCEPT
Microring-based point-to-point silicon photonic link linking two processor/memory nodes

Depiction of a multidie packaging solution based on direct bonding of both chips on a shared substrate or silicon carrier
(a) The optical die would contain the driver and receiver electronics to the extent needed in close proximity to the optical devices. In-package electrical wires transfer data to and from the optical die. Microring-based point-to-point silicon photonic link
(b) Two optical dies communicate over fiber. Each die includes a transmit module based on a microring modulator array and a receive module based on a two-stage microring demultiplexing array. Germanium photodetectors provide feedback for thermal stabilization and high-speed signal detection. Laser sources are assumed to be off-chip separate units
<!-- Slide number: 23 -->
I/O Bandwidth Density

Optical losses increase with narrower channel spacing
More bandwidth can be achieved at the cost of increased laser power
The 25 Gbps channels are heavily penalized by the lower detector sensitivity
Outcome: modulating at a lower rate pays off in terms of aggregate bandwidth
<!-- Slide number: 24 -->
I/O Power Efficiency

With a 1% wall-plug efficiency of current, individually packaged DFB laser technology, the link is infeasible
With a 10% efficiency, the 12.5 Gbps link yields 2.5 pJ/bit
1 pJ/bit could (and should) be achieved by
Tighter integration between photonics and CMOS electronics
relaxing channel spacing (but then you lose bandwidth)
How to make this happen? Improved laser efficiencies, and receiver sensitivities!
<!-- Slide number: 25 -->
Vision of Photonic NoC Integration(that is, optical networks-on-chip)

photonic NoC
3D memory
layers
multi-core
processor layer
Columbia University
Notes:
<!-- Slide number: 26 -->
Optical switching – active networks
bar
in0
out0
PUMPING
(only for active networks)
cross
in1
out1
Optical paths can be established by tuning the resonant wavelength of microring resonators (one wavelength enough for transmission to different directions)
Ring Free Spectral Range
Resonance wavelength
Transmission
This is the principle of broadband photonic switching
<!-- Slide number: 27 -->
Optical switching – passive networks
bar
NO
PUMPING
(passive networks)
in0
out0
cross
in1
out1
Optical paths are univocally associated to the wavelength of the
optical signal (one wavelength NOT enough for transmission to different directions)
Transmission
Broadband photonic switching is simple on this switching structure,
although not that simple on the topologies built around this principle
<!-- Slide number: 28 -->
Optical Components for Designing Optical on-chip Buses

coupler for
attaching fiber to on-chip waveguide
transmitter including driver and ring
modulator for λ1
multiple transmitters including drivers and ring modulators
for each of
λ 1- λ4
receiver including passive ring filter for λ1 and
Photo-detector
receiver including active ring filter for λ1
and photo-detector
passive ring filter for λ1
active ring filter for λ1
<!-- Slide number: 29 -->
Kinds of Optical on-chip Buses
Publicité
SWBR (Single Writer Broadcast Reader)
SWMR (Single Writer Multiple Reader)
MWSR (Multiple Writer Single Reader)
MWMR (Multiple Writer Multiple Reader)
<!-- Slide number: 30 -->
SWBR (Single Writer Broadcast Reader)

A single input terminal modulates the bus wavelength that is then broadcast to all four output terminals.
An SWBR bus requires significant optical power to broadcast packets.
Not very common
Only one wavelength is used.
Notes:
La potenza alta perche per arrivare all’O1 con una potenza rilevabile con il detectore allora dobbiamo moltiplicare la potenza per 4 a parte la perdite durante il passaggio perche 3 frazioni vengono assorbite da altri 3 output precedenti .
<!-- Slide number: 31 -->
SWMR (Single Writer Multiple Reader)
All ring filters are detuned by default , except the one that is assigned to receive the packet from I1, which is actively tuned into the bus wavelength.
I1 has to ensure the activation of the destination ring by logic control that requires additional optical or electrical communication .
Only one wavelength is used.

Can we make it fully passive?
I1 should use a dedicated wavelength for each destination!
with 4 wavelengths no selection logic is needed....
....and communications to more destinations in parallel are potentially feasible
(provided the initiator is able to do that)
4 modulators are needed at the initiator!


…..
<!-- Slide number: 32 -->
Alternative Solutino for Passive SWMR
Exploits the Spatial Division Multiplexing principle








Trade-off between number of laser sources
and number of optical waveguides
<!-- Slide number: 33 -->
MWSR (Multiple Writer Single Reader)
There are four input terminals arbitrating to
modulate the bus wavelength, which is then
dropped at a single output terminal.
MWSR buses require global arbitration, which can be implemented either electrically or optically. (e.g. token ring arbitration in optics).
Only one wavelength is used.

Can we avoid global arbitration (that is, achieve contention-free communication)?
I1-4 should use dedicated wavelengths for communications to O1!
with 4 wavelengths no arbitration logic is needed....
....and communications from multiple initiators in parallel are potentially feasible
the read-out circuit of O1 should be replicated!


.....
Alternatively: spatial-division multiplexing
<!-- Slide number: 34 -->
MWMR (Multiple Writer Multiple Reader)
There are four input terminals arbitrating to modulate the bus wavelength, which is then dropped at one of the four outputs that is actively tuned to the bus wavelength.
The granted input has to ensure the activation of the destination ring by means of control logic that requires additional optical or electrical communication.
Only one wavelength is used

MWMR Bus
Can we make this passive and conflict-free?
4x4=16 wavelengths would be needed
a way too much!
Notes:
By combining nanophotonic SWBR and MWSR buses it is possible to implement : command , write-data, and read-data buses in a DRAM memory channel.
<!-- Slide number: 35 -->
Wavelength-Routed
Let us play with the other degrees of freedom:
Topology, core positioning, SDM, orientation!
O4
O4
I4
I4
I1
I1
O3
O3
orientation
orientation
O1
I3
O1
I3
I2
O2
I2
O2
Connectivity is feasible with 3 wavelengths!

| | O1 | O2 | O3 | O4 |
| --- | --- | --- | --- | --- |
| I1 | wvl1 | wvl2 | wvl2 | wvl1 |
| I2 | wvl1 | wvl1 | wvl3 | wvl3 |
| I3 | wvl2 | wvl1 | wvl1 | wvl2 |
| I4 | wvl3 | wvl3 | wvl1 | wvl1 |
wvl1
wvl2
wvl3
This has evolved into a non-blocking (wavelength-routed) CROSSBAR!
Notes:
<!-- Slide number: 36 -->
WAVELENGTH-ROUTING
WAVELENGTH-ROUTING
ALL-TO-ALL CONNECTIVITY
CONTENTION-FREE COMMUNICATION
ALL-TO-ALL COMMUNICATIONS ARE POTENTIALLY FEASIBLE AT THE SAME TIME WITHOUT ANY CONFLICT
THE ROUTING PATH IS UNIVOCALLY DEFINED BY THE WAVELENGTH CHOSEN FOR COMMUNICTION, UNIVOCALLY ASSOCIATED WITH ONE RECEIVER.
COSTS A LOT OF WAVELENGTHS!
THE COST CAN BE AMORTIZED THROUGH SPATIAL DIVISION MULTIPLEXING
<!-- Slide number: 37 -->
Nanophotonic crossbars
Nanophotonic crossbars use a dedicated nanophotonic bus per terminal to enable every input terminal to send a packet to a different output terminal
at the same time.
We illustrate 5 types:
SWMR Crossbar
MWSR Crossbar
MWMR Crossbar
Wavelength-routed Crossbar
Space-routed Crossbar
<!-- Slide number: 38 -->
SWMR Crossbar
There is one bus per input and every output can read from any bus.
SWMR crossbars usually include a low bandwidth SWBR crossbar to implement distributed redundant arbitration at the output terminals and/or to determine which receivers at the destination should be actively tuned.
A buffered SWMR crossbar avoids the need for any global or distributed arbitration.

Notes:
SWBR crossbars are also possible where the packet is broadcast to all output terminals, and each output terminal is responsible for converting the packet into the electrical domain and determining if the packet is actually destined for that terminal.
As an example, if I2 wants to send a packet to O3 it first arbitrates for access to the output terminal, then (assuming it wins arbitration) the receiver for wavelength λ2 at O3 is actively tuned.
<!-- Slide number: 39 -->
SWMR crossbar variants
Naive passivation does not yield conflict-freedom
wvl1 from I1
wvl1 from I2
wvl1 from I3
wvl1 from I4

O1 receives:
.....
.....
wvl4 from I1
wvl4 from I2
Wvl4 from I3
Wvl4 from I4
O4 receives:
Decoding is solved, but distributed arbitration is still needed, and cannot be avoided with buffered outputs due to interferences! Also, this has evolved into a MWSR crossbar!
Spatial Division Multiplexing

...........
Decoding as well as global arbitration removed, as long as output buffering is there!
<!-- Slide number: 40 -->
MWSR Crossbar
Uses one bus per output and allows every input to write any of these buses.
Implements distributed arbitration between the input terminals.

Spatial Division Multiplexing

With buffered outputs, no arbitration needed
Notes:
As an example, if I2 wants to send a packet to O3 it first arbitrates, and then (assuming it wins arbitration) it modulates wavelength λ3.
<!-- Slide number: 41 -->
MWMR Crossbar

Arbitartion at the trasmission side is required to get access to a given wavelength.
Arbitration is required to compete for a destination, unless buffering is implemented.
Notes:
There have been several diverse proposals for implementing global crossbars with nanophotonics such as SWBR,MRBR,
<!-- Slide number: 42 -->
Wavelength Routing

Arbitrary
interconnection
topology
Publicité
......
......
Usage model
λ1
O1
I1
λ4
λ2
λ3
O2
λ3
λ4
O3
λ1
λ2
O4
I4
BENEFITS
No time is spent in Routing/Decoding and Arbitration.
High communication performance predictability
......
CHALLENGE: HARD TO SCALE TO A LARGE NUMBER OF COMMUNICATION ACTORS
<!-- Slide number: 43 -->
A wavelength-routed topology
Topology differentiator: enable wavelength-routing with the lowest amount of resources:
laser sources, no. and kind of microring resonators.

Pse 2x2


Each (couple of) microring(s) has its own radius to enable resonance on a specific wavelength
<!-- Slide number: 44 -->
Broadband Switching in wavelength-routed topologies
Broadband passive switching consists of embedding Multiple Virtual Networks into the same set of waveguides
One possibility is to leverage as much as possible the wavelengths in the resonance band of the Micro Ring Resonators (MRR) of the NoC’s add-drop filters.
Transmission Responses with different values of radius

R1
R2
Cascaded 1x2PSEs
λ2
N
λ2
λ1
λ2
E
W
λ1
S
λ1
10 Gbit/s
10 Gbit/s
Cascaded 1x2PSEs
λ2,1
λ2
λ2,1
λ2
N
λ1,1
λ1
λ2
E
W
λ1
S
λ1,1
λ1
20 Gbit/s
20 Gbit/s
λ2
λ1
λ2,1
λ1,1
Transmission
Wavelength
This overlapping provides routing faults. Radius design and wavelength selection should be carefully engineered. In any case, high levels of bit parallelism are not feasible!
<!-- Slide number: 45 -->
Wavelength-Routed Topologies
I8
I7
I6
I5
I4
I3
I2
I1
T1
T4
T7
T2
T3
T5
T6
T8


8x8 GWOR
8x8 λ-Router
8x8 Folded Crossbar
Multi-stage connectivity pattern vs.
More or less (MRR) populated grid-like structures vs.
Wrap-around extensions
| TOPOLOGY | Total # of MRRs | MAX # of Crossings Logic Scheme |
| --- | --- | --- |
| 8x8 λ-Router | 56 | 7 |
| 8x8 GWOR | 48 | 10 |
| 8x8 Folded Crossbar | 64 | 14 |
Min(no. of hops)
Higher no. of hops in grids
More extended grid

<!-- Slide number: 46 -->
Placement and Routing constraints
Target Platform for chip-scale optical interconnect technology:
3D stacking of processing, memory and optical layers
Placement Constraints: It is reasonable to assume that the Hubs are positioned in the middle of the clusters

Placement Constraints: The Memory Controllers are positioned pairwise
at opposite sides of the chip thus reflecting industrial practice (e.g., TILE64)
Notes:
<!-- Slide number: 47 -->
The Design Predictability Gap






8x8 λ-Router Real Layout
8x8 Folded Crossbar Real Layout
8x8 GWOR Real Layout
Layout of the 8x8 Folded Crossbar is much more regular than that of the 8x8 λ-Router and the 8x8 GWOR due to a ring-like structure





THE WORST LOGIC TOPOLOGIES MAY BECOME THE BEST PHYSICAL ONES!
<!-- Slide number: 48 -->
How to make Compelling Cases
Joshi et al. and Ramini et al. started to put together some recommendations for trustworthy assessement of optical interconnect technology
#.1 Clearly specify the Logical Topology
#.2 Explore the Space of Mapping options to nanophotonic devices
#.3 Account for Place&Route constraints
#.4 Keep it simple to minimize risk
#.5 Consider the network interface architecture
#.6 Use an aggressive electrical baseline
#.7 Assume a broad range of device parameters
#.8 Carefully consider static power overheads
Trustworthy ONoC vs. ENoC crossbenchmarking
(GOAL OF THIS WORK)
<!-- Slide number: 49 -->
RULE #.1: OUR CHOICE
Let us consider a Tilera-style 4x4 CMP
1,2 GHz operating speed, 32 bit parallelism
Directory-based implementation of the MOESI cache coherence protocol


Messaging from cache coherent protocols is not only throughput-critical, but also latency-critical hence we opted for wavelength-routed xbars
Notes:
We conduct a post-layout charecterization of a 4x4 Mesh topology based on the : 1 Ghz ,….
<!-- Slide number: 50 -->
RULE #.2 and #.3:
Explore the Space of Mapping options to nanophotonic devices and Account for Place&Route constraints
What kind of topology should be selected?
The Optical Ring features a higher degree of design predictability
than multi-stage networks in the presence of place&route constraints
Logic Scheme
Physical Layout
16x16 Optical Ring


VS.
16x16 Multi-Stage ONoC
16x16 Optical Ring
16x16 Multi-Stage ONoC

Publicité
a
r
b
q
c
p
VS.
d
o
e
n
m
f
g
l
h
i
The Optical Ring exhibits x3 lower Insertion loss (ILmax) than Multi-Stage ONoC
<!-- Slide number: 51 -->
RULE #.4: Keep it simple to minimize the risk (Principle of the Optical Ring)
The same wavelengths can be reused on a single waveguide
to establish multiple and non-overlapping communications

Inspired by S.Le Beux, DATE 2011
Two wavelengths (λ1 , λ2) and two physical waveguides (blue and black) are sufficient to realize 12 contention-free optical paths
Notes:
By adding further wavelengths (laser sources) or by replicating different physical waveguides it is possible to achieve an arbitrary level of bit parallelism
<!-- Slide number: 52 -->
RULE #.2 and #.3:
RULE #.3 : Optical Ring with Physical Awareness
Waveguide Crossings Concern in Optical Ring Topology


IN ORDER TO ACCURATELY DESIGN AN OPTICAL RING TOPOLOGY, WAVEGUIDE CROSSINGS CANNOT BE OVERLOOKED!!!!!!!

MICRO-RING-RESONATOR OVERHEAD
NOT ONLY RING FILTERS AT THE DESTINATION STAGE SHOULD BE ACCOUNTED FOR, BUT ALSO COUPLERS INTO THE RING WAVEGUIDES
<!-- Slide number: 53 -->
RULE #.2,3 and 4 : 16X16 Competing Optical Ring
The 16x16 Optical Ring is vertically stacked on top
of an electronic layer composed of 16 Network Interfaces
Layout of a 16x16 Ring Topology

3D STACKING


1-bit parallelism: 13 wavelengths reused across 16 different waveguides
to enable 240 contention-free optical paths.
Explored parallelism: from 3 to 4 bits
<!-- Slide number: 54 -->
RULE #.2,3 and 4 : Laser Power Assessment
For the sake of a more comprehensive analysis of nanophotonic devices
we distinguished two relevant cases : REALISTIC & AGGRESSIVE

Laser Efficiency=20%
MMI Taper
(Multi-Mode Interference)
Laser Efficiency=8%
Elliptical Taper
THERE IS AN INSERTION LOSS CRITICAL PATH FOR EACH WAVELENGTH
SOME LASER SOURCES ARE MORE POWER HUNGRY THAN OTHERS


Vs.
IL=0,52 dB
IL=0,18 dB
WITH AN AGGRESSIVE TECHNOLOGY, THE POWER GAP IS STRONGLY REDUCED (4x). HOWEVER, THE RELATIVE GAP ACROSS WAVELENGTHS STILL PERSISTS
LASER SOURCES (ASSUMED TO BE CONTINUOUS WAVE) MUST BE TREATED IN A DIFFERENT WAY
Notes:
<!-- Slide number: 55 -->
RULE #.5: CONSIDER THE NETWORK INTERFACE
3 inj. Buffers vs. MESSAGE DEPENDENT DEADLOCK
3 serializers per destination
(3-bit parallelism)

Credit-based
flow control
Source-Synchronous Clock
Mesochronous synchronizers
FIFO size covering round trip delay
<!-- Slide number: 56 -->
RULE #.6: USE AN AGGRESSIVE ELECTRICAL BASELINE
Let us consider a Tilera-style 4x4 CMP
1,2 GHz operating speed, 32 bit parallelism
Consolidated xpipesLite NoC architecture
(simplest design point for High-end Embedded Systems)
3 virtual channels to avoid message dependent deadlock

Multi-Switch Approach (Low power Approach)
F. Gilabert, M.E. Gomez, S. Medardoni, D. Bertozzi.
“ Improved utilization of noc channel bandwidth by switch replication for costeffectivemultiprocessorsystems-on-chip”. (NOCS), 2010.

Real synthesis runs on an
industrial low-power SVT 40nm technology library!
Clock gating applied!
Notes:
We conduct a post-layout charecterization of a 4x4 Mesh topology based on the : 1 Ghz ,….
<!-- Slide number: 57 -->
RULE #.7: USE A BROAD RANGE OF DEVICE PARAMETERS
RULE #.8: CAREFULLY CONSIDER STATIC POWER
CONSERVATIVE PARAMETERS
AGGRESSIVE PARAMETERS
WALL-PLUG LASER EFFICIENCY=8%
CROSSING OPTIMIZATION= ELLIPTICAL TAPER
THERMAL TUNING=20μW/ring
TRANSMITTER DYNAMIC ENERGY =50fj/bit
TRANSMITTER FIXED ENERGY =10 fj/bit
RECEIVER DYNAMIC ENERGY =25fj/bit
RECEIVER FIXED ENERGY=15fj/bit
WALL-PLUG LASER EFFICIENCY=20%
CROSSING OPTIMIZATION= MMI TAPER
THERMAL TUNING=20μW/ring
TRANSMITTER DYNAMIC ENERGY =20fj/bit
TRANSMITTER FIXED ENERGY=2,5 fj/bit
RECEIVER DYNAMIC ENERGY =10fj/bit
RECEIVER FIXED ENERGY=5 fj/bit
A CONSISTENT SET OF STATIC VS. DYNAMIC POWER CONTRIBUTIONS
HAVE BEEN SELECTED FROM THE SAME SOURCE IN THE OPEN LITERATURE
S. Beamer, C. Sun, Y. Kwon, A. Joshi, C. Batten, V. Stojanovic, K.Asanovic.
Re-architecting DRAM memory systems with monolithically integrated silicon photonics. ISCA ’10
<!-- Slide number: 58 -->
Experimental Results
<!-- Slide number: 59 -->
Performance Speed-up
GEM5 @ PARSEC2.1 BENCHMARK SUITE

+18% @ 3bit
+23% @ 4bit
More than 4 bits parallelism would incur too much static power overhead.
Less than 3 bits would provide a lower link bitrate than the electronic counterpart.
ONoC technology can actually deliver significant performance speedup!
Even in cache coherent systems, which are not communication-intensive.
Upper bound: dictated by injection frequency of electronics into interface dc-FIFO
(4 bit-parallelism is close to saturation)
Notes:
<!-- Slide number: 60 -->
Realistic Modeling Framework
THE XBENCH. RULES ARE APPLIED
CONSERVATIVE OPTICAL TECHNOLOGY

ONoC is studied @3,4 bit parallelism
A conservative optical technology is still far away from the break-even point
The role of the optical network interface now comes to the forefront
Static energy of ONoC (plus its network interfaces) is impressive!
<!-- Slide number: 61 -->
Realistic Modeling Framework
AGGRESSIVE OPTICAL TECHNOLOGY


ONoC is studied @3,4 bit parallelism
An aggressive optical technology gets closer to the break-even point
(about 11,6%, @3bit) WITHOUT ACHIEVING IT!
The static energy cost of the optical network interfaces is higher than
(almost 2 times@ 3bit ) that of the ONoC.
<!-- Slide number: 62 -->
Didn’t we overlook something?
A good interconnect fabric implies that it is a small contributor to total system energy
An interconnect fabric speeding up application execution causes the system to burn less energy
ASSUMPTIONS

The whole System burns 15 Watts.
3 and 4bit parallelism for ONoC.
Both AGGRESSIVE AND CONSERVATIVE technologies

Normalized System Energy
Chart
| Category | 3 bit | 4 bit |
|---|---|---|
| ENoC | 1.0 | 1.0 |
| ONoC (Conservative) | 0.8475 | 0.7877000000000005 |
| ONoC (Aggressive) | 0.8225 | 0.7589000000000007 |

21%
24%
ONoC makes the system more energy efficient even with a conservative technology of optical components
<!-- Slide number: 63 -->
Conclusions
Publicité
This secti...