The Adam optimiser updates a parameter with mt=β1mt−1+(1−β1)gt, vt=β2vt−1+(1−β2)gt2, bias-corrected m^t=mt/(1−β1t), v^t=vt/(1−β2t) and step θt=θt−1−ηm^t/(v^t+ϵ), with m0=v0=0. Take β1=0.9, β2=0.999, learning rate η=0.01, ϵ=0 and a first gradient g1=3.7. What is the magnitude of the very first parameter update ∣θ1−θ0∣?