Stochastic Gradient Push for Distributed Deep Learning

Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, Michael Rabbat

MatrixRepresentation.

LABEL:SGPalg,SGPwaspresentedfromnodei′sperspective(foralli∈[n]).However,wecanactuallywritetheSGPupdateateachiterationfromaglobalviewpoint.Toseethis,firstdefinethefollowingmatrices,forallr=1,2,…τ,

all(virtualandnon-virtual)nodes′parametersatiterationk.Recallthattheweinitializeallvirtualnodeswithparametersx(k)=0andpush-sumweightw(0)=0.Additionally,sincethevirtualnodesareonlyusedtomodeldelays,anddonotcomputeanygradientupdates,weusetheconventionthatz(k)=0,ξ(k)=0,and∇F(z(k);ξ(k))=0forallvirtualnodesatalltimesk.Therefore,wedefinetheaugmentedde-biasedparametermatrixandstochastic-seedmatrixasfollows

LABEL:SGPalg(lines19to24inOSGPAlgorithmLABEL:alg:osgp)canbeexpressedfromaglobalperspectiveasfollows