SGD is the perfect algorithm for use in online learning. Except it has one major drawback – is sensitive to feature scaling.
In some of my trials with the SGD learner in scikit-learn, I have seen terrible performance if I don’t do feature scaling.
Which begs the question – How does VW do feature scaling ? After all VW does online learning.
It seems VW uses a kind of SGD that is scale variant: