Deep Learning with Coherent Nanophotonic Circuits

Yichen Shen, Nicholas C. Harris, Scott Skirlo, Mihika Prabhu, Tom Baehr-Jones, Michael Hochberg, Xin Sun, Shijie Zhao, Hugo Larochelle, Dirk Englund, Marin Soljacic

I Optical Neural Network Device Architecture

An artificial neural network (ANN) lecun2015 consists of a set of input artificial neurons (represented as circles in Fig. 1(a)) connected to at least one hidden layer and an output layer. In each layer (depicted in Fig. 1(b)), information propagates by linear combination (e.g. matrix multiplication) followed by the application of a nonlinear activation function. ANNs can be trained by feeding training data into the input layer and then computing the output by forward propagation; weighting parameters in each matrix are subsequently optimized using back propagation schmidhuber2015deep .

The Optical Neural Network (ONN) architecture is depicted in Fig. 1 (b,c). As shown in Fig. 1(c), signals are encoded in the amplitude of optical pulses propagating in integrated photonic waveguides where they pass through an optical interference unit (OIU) and finally an optical nonlinearity unit (ONU). Optical matrix multiplication is implemented with an OIU and nonlinear activation is realized with an ONU.

To realize an OIU that can implement any real-valued matrix, we use the singular value decomposition (SVD) lawson1995solving since a general, real-valued matrix (M\mathbf{M}) may be decomposed as M=UΣV∗\mathbf{M}=\mathbf{U}\bm{\Sigma}\mathbf{V}^{*}, where UU is an m×mm\times m unitary matrix, Σ\Sigma is a m×nm\times n diagonal matrix with non-negative real numbers on the diagonal, and V∗V^{*} is the complex conjugate of the n×nn\times n unitary matrix VV. It was theoretically shown that any unitary transformations U,V∗\mathbf{U},\mathbf{V^{*}} can be implemented with optical beamsplitters and phase shifters PhysRevLett.73.58 ; Miller:15 . Matrix multiplication implemented in this manner consumes, in principle, no power. The fact that a major part of ANN calculations involves matrix products enables the extreme energy efficiency of the ONN architecture presented here. Finally, Σ\bm{\Sigma} can be implemented using optical attenuators; optical amplification materials such as semiconductors or dyes could also be used connelly2007semiconductor .

The ONU can be implemented using optical nonlinearities such as saturable absorption selden1967pulse ; Bao2010 ; schirmer1997nonlinear and bistability soljavcic2002optimal ; PhysRevLett.71.3959 ; PhysRevB.62.R7683 ; Nozaki2010 ; Ros2015 that have been demonstrated seperately in photonic circuits. For an input intensity IinI_{in}, the optical output intensity is thus given by a nonlinear function Iout=f(Iin)I_{out}=f(I_{in}) NIPS2012_4824 .

II Experiment

For an experimental demonstration of our ONN architecture, we implement a two layer, fully connected neural network with the OIU shown in Fig. 2 and use it to perform vowel recognition. To prepare the training and testing dataset, we use 360 datapoints that each consist of four log area ratio coefficients chow2004speaker of one phoneme. The log area ratio coefficients, or feature vectors, represent the power contained in different logarithmically-spaced frequency bands and are derived by computing the Fourier transform of the voice signal multiplied by a Hamming window function. The 360 datapoints were generated by 90 different people speaking 4 different vowel phonemes deterding1990speaker . We use half of these datapoints for training and the remaining half to test the performance of the trained ONN. We train the matrix parameters used in the ONN with the standard back propagation algorithm using stochastic gradient descent method hinton2006reducing , on a conventional computer. Further details on the dataset and backpropagation procedure are included in Supplemental Information Section 3.

The coherent ONN is realized with a programmable nanophotonic processor Harris:2015ux composed of an array of 56 Mach-Zehnder interferometers (MZIs) and 213 phase shifting elements, as shown in Fig. 2. Each interferometer is composed of two evanescent-mode waveguide couplers sandwiching an internal thermo-optic phase shifter harris2014efficient to control the splitting ratio of the output modes, followed by a second modulator to control the relative phase of the output modes. By controlling the phase imparted by these two phase shifters, these MZIs perform all rotations in the SU(2)SU(2) Lie group given a controlled incident phase on the two electromagnetic input modes of the MZI. The nanophotonic processor was fabricated in a silicon-on-insulator photonics platform with the OPSIS Foundry baehr201225 .

To experimentally realize arbitrary matrices by SVD, we programmed an SU(4)SU(4) core PhysRevLett.73.58 ; miller2013selfconfiguring and a non-unitary diagonal matrix multiplication core (DMMC) into the nanophotonic processor Harris:2015ux ; harris2014efficient , as shown in Fig. 2 (b). The SU(4)SU(4) core implements operators UU and VV by a Givens rotations algorithm PhysRevLett.73.58 ; miller2013selfconfiguring that decomposes unitary matrices into sets of phase shifters and beam splitters, while the DMMC implements Σ\Sigma by controlling the splitting ratios of the DMMC interferometers to add or remove light from the optical mode relative to a baseline amplitude. The measured fidelity for the 720 OIU and DMMC cores used in the experiment was 99.8 ±\pm 0.003 %; see methods for further detail.

In this analog computer, fidelity is limited by practical non-idealities such as (1) finite precision with which an optical phase can be set using our custom 240-channel voltage supply with 16-bit voltage resolution per channel (2) photodetection noise, and (3) thermal cross-talk between phase shifters which effectively reduces the number of bits of resolution for setting phases. As with digital floating-point computations, values are represented to some number of bits of precision, the finite dynamic range and noise in the optical intensities causes effective truncation errors. A detailed analysis of finite precision and low-flux photon shot noise is presented in Supplement Section 1.

In this proof-of-concept demonstration, we implement the nonlinear transformation Iout=f(Iin)I_{out}=f(I_{in}) in the electronic domain, by measuring optical mode output intensities on a photodetector array and injecting signals IoutI_{out} into the next stage of OIU. Here, ff models the mathematical function associated with a realistic saturable absorber (such as a dye, semiconductor or graphene saturable absorber or saturable amplifier) that could, in future implementations, be directly integrated into waveguides after each OIU stage of the circuit. For example, graphene layers integrated on nanophotonic waveguides have already been demonstrated as saturable absorbers cheng2014plane . Saturable absorption is modeled as selden1967pulse (Supplement Section 2),

where σ\sigma is the absorption cross section, τs\tau_{s} is the radiative lifetime of the absorber material, T0T_{0} is the initial transmittance (a constant that only depends on the design of saturable absorbers), I0I_{0} is the incident intensity, and TmT_{m} is the transmittance of the absorber. Given an input intensity I0I_{0}, one can solve for Tm(I0)T_{m}(I_{0}) from Eqn. 1, and the output intensity can be calculated as Iout=I0⋅Tm(I0)I_{out}=I_{0}\cdot T_{m}(I_{0}). A plot of the saturable absorber’s response function Iout(Iin)I_{out}(I_{in}) is shown in supplement Section 2.

After programming the nanophotonic processor to implement our ONN architecture, which consist of 4 layers of OIUs with 4 neurons on each layer (which requires training a total of 4⋅6⋅2=484\cdot 6\cdot 2=48 phase shifter settings), we evaluated it on the vowel recognition test set. Our ONN correctly identified 138/180 cases (76.7%) endnote1 compared to a simulated correctness of 165/180 (91.7%).

Since our ONN processes information in the analog signal domain, the architecture can be vulnerable to computational errors. Photodetection and phase encoding are the dominant sources of error in the ONN presented here (as discussed above). To understand the role of phase encoding noise and photodection noise in our ONN hardware architecture and to develop a model for its accuracy, we numerically simulate the performance of our trained matrices with varying degrees of phase encoding noise (σΦ\sigma_{\Phi}) and photodection noise (σD\sigma_{D}) (detailed simulation steps can be found in methods section). The distribution of correctness percentage vs σΦ\sigma_{\Phi} and σD\sigma_{D} is shown in Fig. 3 (a), which serves as a guide to understanding experimental performance of the ONN. Improvements to the control and readout hardware, including implementing higher precision analog-to-digital converters in the photodetection array and voltage controller, are practical avenues towards approaching the performance of digital computers. Well-known techniques can be applied to engineer the photodiode array to achieve significantly higher dynamic range; for example, using logarithmic or multi-stage gain amplifiers. Addressing these managable engineering problems can further enhance the correctness performance of the ONN to ultimately achieve correctness percentages approaching those of error-corrected digital computers. In addition, ANN parameters trained by conventional back propagation algorithm can become suboptimal when encoding errors are encountered. In such a case, robust simulated annealing algorithms bertsimas2010robust can be used to train ANN parameters which is error-tolerant, hence when encoded in the ONN, will have better performance.

III Discussion

Processing big data at high speeds and with low power is a central challenge in the field of computer science, and, in fact, a majority of the power and processors in data centers are spent on doing forward propagation (test-time prediction). Furthermore, low forward propagation speeds limit applications of ANNs in many fields including self-driving cars which require high speed and parallel image recognition.

Our optical neural network architecture takes advantage of high detection rate, high-sensitivity photon detectors to enable high-speed, energy-efficient neural networks compared to state-of-the-art electronic computer architectures. Once all parameters have been trained and programmed on the nanophotonic processor, forward propagation computing is performed optically on a passive system. In our implementation, maintaining the phase modulator settings requires some (small) power of ∼\sim10 mW per modulator on average. However, in future implementations, the phases could be set with nonvolatile phase-change materialsrios2015integrated , which would require no power to maintain. With this change, the total power consumption is limited only by the physical size, the spectral bandwidth of dispersive components (THz), and the photo-detection rate (100GHz). In principle, such a system can be at least 2 orders of magnitude faster than electronic neural networks (which are restricted at GHz clock rate). Assuming our ONN has N nodes, implementing mm layers of N×\timesN matrix multiplication and operating at a typical 100 GHz photo-detection rate, the number of operations per second of our system would be

ONN power consumption during computation is dominated by the optical power necessary to trigger an optical nonlinearity and achieve a sufficiently high signal-to-noise ratio (SNR) at the photodetectors (assuming shot-noise limited detection on nn photons per pulse, SNR≃1/nSNR\simeq\sqrt{1/n}). We assume a saturable absorber threshold of p≃1p\simeq 1 MW/cm2 – valid for many dyes, semiconductors, and graphene selden1967pulse ; Bao2010 . Since the cross section for the waveguide is A=0.2\upmuA=0.2\upmum ×0.5\upmu\times 0.5\upmum, the total power needed to run the system is therefore estimated to be: P≈N mWP\approx N~{}\text{mW}. Therefore, the energy per operation of ONN will scale as R/P=2m⋅N⋅1014R/P=2m\cdot N\cdot 10^{14} operations/J (or P/R=5mNP/R=\frac{5}{mN} fJ/operation). Almost the same energy performance and speed can be obtained if optical bistability soljavcic2002optimal ; Nozaki2010 ; Tanabe:05 is used instead of saturable absorption as the enabling nonlinear phenomenon. Even for very small neural networks, the above power efficiency is already at least 3 orders of magnitude better than that in conventional electronic CPUs and GPUs, where P/R≈1P/R\approx 1pJ/operation (not including the power spent on data movement) horowitz20141 , while conventional image recognition tasks require tens of millions of training parameters and thousands of neurons (mN≈105mN\approx 10^{5})krizhevsky2012imagenet . These considerations suggest that the optical NN approach may be tens of millions times more efficient than conventional computers for standard problem sizes. In fact, the larger the neural network, the bigger the advantage of using optics is: this comes from the fact that evaluating an N×NN\times N matrix in electronics requires O(N2)O(N^{2}) energy, while in optics, it requires in principle no energy. Further details on power efficiency calculation can be found in the Supplementary information section 3.

ONNs enable new ways to train ANN parameters. On a conventional computer, parameters are trained with back propagation and gradient descent. However, for certain ANNs where the effective number of parameters substantially exceeds the number of distinct parameters (including recurrent neural networks (RNN) and convolutional neural networks(CNN)), training using back propagation is notoriously inefficient. Specifically the recurrent nature of RNNs gives them effectively an extremely deep ANN (depth=sequence length), while in CNNs the same weight parameters are used repeatedly in different parts of an image for extracting features. Here we propose an alternative approach to directly obtain the gradient of each distinct parameter without back propagation, using forward propagation on ONN and the finite difference method. It is well known that the gradient for a particular distinct weight parameter ΔWij\Delta W_{ij} in ANN can be obtained with two forward propagation steps that compute J(Wij)J(W_{ij}) and J(Wij+δij)J(W_{ij}+\delta_{ij}), followed by the evaluation of ΔWij=J(Wij+δij)−J(Wij)δij\Delta W_{ij}=\frac{J(W_{ij}+\delta_{ij})-J(W_{ij})}{\delta_{ij}} (this step only takes two operations). On a conventional computer, this scheme is not favored because forward propagation (evaluating J(W)J(W)) is computationally expensive. In an ONN, each forward propagation step is computed in constant time (limited by the photodetection rate which can exceed 100 GHz Vivien:12 ), with power consumption that is only proportional to the number of neurons–making the scheme above tractable. Furthermore, with this on-chip training scheme, one can readily parametrize and train unitary matrices–an approach known to be particularly useful for deep neural networks arjovsky2015unitary . As a proof of concept, we carry out the unitary-matrix-on-chip training scheme for our vowel recognition problem (see Supplementary Information Section 4).

Regarding the physical size of the proposed ONN, current technologies are capable of realizing ONNs exceeding the 1000 neuron regime – photonic circuits with up to 4096 optical components have been demonstrated Sun2013 . 3-D photonic integration could enable even larger ONNs by adding another spatial degree of freedom rechtsman2013photonic . Furthermore, by feeding in input signals (e.g. an image) via multiple patches over time (instead of all at once) – an algorithm that has been increasingly adopted by deep learning community Jia:Caffe2014 – the ONN should be able to realize much bigger effective neural networks with relatively small number of physical neurons.

IV Conclusion

The proposed architecture could be applied to other artificial neural network algorithms where matrix multiplications and nonlinear activations are heavily used, including convolutional neural networks and recurrent neural networks. Further, the superior forward propagation speed and power efficiency of our ONN can potentially enable training the neural network on the photonics chip directly, using only forward propagation. Finally, it needs to be emphasized that another major portion of power dissipation in current NN architectures is associated with data movement–an outstanding challenge that remains to be addressed. However, recent dramatic improvements in optical interconnects using integrated photonics technology has the potential to significantly reduce data-movement energy cost sun2015nature . Further integration of optical interconnects and optical computing units need to be explored to realize the full advantage of all-optical computing.

V Methods

Fidelity Analysis We evaluated the performance of the SU(4)SU(4) core with the fidelity metric f=∑ipiqif=\sum_{i}\sqrt{p_{i}q_{i}} where pi,qip_{i},q_{i} are experimental and simulated normalized (∑ixi=1\sum_{i}x_{i}=1 where x∈{p,q}x\in\{p,q\}) optical intensity distributions across the waveguide modes, respectively.

Simulation Method for Noise in ONN We carry out the following steps to numerically simulate the performance of our trained matrices with varying degrees of phase encoding (σΦ\sigma_{\Phi}) and detection (σD\sigma_{D}) noise.

For each of the four trained 4×44\times 4 unitary matrices UkU^{k}, we calculate a set of {θik,ϕik}\{\theta_{i}^{k},\phi_{i}^{k}\} that encode the matrix.

We add a set of random phase encoding errors, {δθik,δϕik}\{\delta\theta_{i}^{k},\delta\phi_{i}^{k}\} to the old calculated phases {θik,ϕik}\{\theta_{i}^{k},\phi_{i}^{k}\}, where we assume each δθik\delta\theta_{i}^{k} and δϕik\delta\phi_{i}^{k} is a random variable sampled from a Gaussian distribution G(μ,σ)G(\mu,\sigma) with μ=0\mu=0 and σ=σΦ\sigma=\sigma_{\Phi}. We obtain a new set of perturbed phases {θik′,ϕik′}={θik+δθik,ϕik+δϕik}\{{\theta_{i}^{k}}^{\prime},{\phi_{i}^{k}}^{\prime}\}=\{\theta_{i}^{k}+\delta\theta_{i}^{k},\phi_{i}^{k}+\delta\phi_{i}^{k}\}.

We encode the four perturbed 4×44\times 4 unitary matrices Uk′{U^{k}}^{\prime} based on the new perturbed phases {θik′,ϕik′}\{{\theta_{i}^{k}}^{\prime},{\phi_{i}^{k}}^{\prime}\}.

We carry out the forward propagation algorithm based on the perturbed matrices Uk′{U^{k}}^{\prime} with our test data set. During the forward propagation, every time when a matrix multiplication is performed (let’s say when we compute v→=Uk′⋅u→\overrightarrow{v}={U^{k}}^{\prime}\cdot\overrightarrow{u}), we add a set of random photo-detection errors δv→\overrightarrow{\delta v} to the resulting v→\overrightarrow{v}, where we assume each entry of δv→\overrightarrow{\delta v} is a random variable sampled from a Gaussian distribution G(μ,σ)G(\mu,\sigma) with μ=0\mu=0 and σ=σD⋅∣v→∣\sigma=\sigma_{D}\cdot|\overrightarrow{v}|. We obtain the perturbed output vector v→′=v→+δv→\overrightarrow{v}^{\prime}=\overrightarrow{v}+\overrightarrow{\delta v}.

With the modified forward propagation scheme above, we calculate the correctness percentage for the perturbed ONN.

Steps 2)-5) are repeated 50 times to obtain the distribution of correctness percentage for each phase encoding noise (σΦ\sigma_{\Phi}) and photodetection noise (σD\sigma_{D}).

VI Acknowledgements

We thank Yann LeCun, Max Tegmark, Isaac Chuang and Vivienne Sze for valuable discussions. This work was supported in part by the Army Research Office through the Institute for Soldier Nanotechnologies under contract W911NF-13-D0001, and in part by the Air Force Office of Scientific Research Multidisciplinary University Research Initiative (FA9550-14-1-0052) and the Air Force Research Laboratory RITA program (FA8750-14-2-0120). M.H. acknowledges support from AFOSR STTR grants, numbers FA9550-12-C-0079 and FA9550-12-C-0038 and Gernot Pomrenke, of AFOSR, for his support of the OpSIS effort, though both a PECASE award (FA9550- 13-1-0027) and funding for OpSIS (FA9550-10-1-0439). N. H. acknowledges support from the National Science Foundation Graduate Research Fellowship grant no. 1122374.

VII Author Contributions

Y.S., N.H., S.S., X.S., S.Z., D.E., and M.S., developed the theoretical model for the optical neural network. N.H., M.P., and Y.S. performed the experiment. N.H. developed the cloud-based simulation software for collaborations with the programmable nanophotonic processor. Y.S., S.S., and X.S. prepared the data and developed the code for training MZI parameters. T.B.J. and M.H. fabricated the photonic integrated circuit. All authors contributed in writing the paper.

References