Stochastic Gradient Push for Distributed Deep Learning
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, Michael Rabbat
MatrixRepresentation.
LABEL:SGPalg,SGPwaspresentedfromnodei′sperspective(foralli∈[n]).However,wecanactuallywritetheSGPupdateateachiterationfromaglobalviewpoint.Toseethis,firstdefinethefollowingmatrices,forallr=1,2,…τ,
all(virtualandnon-virtual)nodes′parametersatiterationk.Recallthattheweinitializeallvirtualnodeswithparametersx(k)=0andpush-sumweightw(0)=0.Additionally,sincethevirtualnodesareonlyusedtomodeldelays,anddonotcomputeanygradientupdates,weusetheconventionthatz(k)=0,ξ(k)=0,and∇F(z(k);ξ(k))=0forallvirtualnodesatalltimesk.Therefore,wedefinetheaugmentedde-biasedparametermatrixandstochastic-seedmatrixasfollows
LABEL:SGPalg(lines19to24inOSGPAlgorithmLABEL:alg:osgp)canbeexpressedfromaglobalperspectiveasfollows