View on GitHub

Word2vec-fork

Extensions to Tomas's word2vec

Download this project as a .zip file Download this project as a tar.gz file

Extensions to Tomas's word2vec

This project makes some extension to the word2vec tool.

Binary Tree for HS using word embedding info

Rebuild the binary tree every several iterations, using the temporary word embedding vectors. Get better performance than the original HS, with a little overhead of speed.

Benchmark

Tests running on a 12-core CPU, parameters are all the same in the original scripts, except the '-threads' is set to 12. Baselines are using the original word2vec tool.

Word accuracy

Using command:

./word2vec -train text8 -output vectors.bin -cbow 1 -size 200 -window 8 -negative 25 -hs 0 -sample 1e-4 -threads 12 -binary 1 -iter 15
./compute-accuracy vectors.bin 30000 < questions-words.txt
Total accuracy(%) Semantic accuracy(%) Syntactic accuracy(%) Speed(kwords/s)
Neg-25 52.74 57.35 50.42 98.49
Neg-5 45.22 43.04 46.31 322.00
HS 34.39 35.22 33.97 186.09
HS-rebuild 42.97 47.67 40.60 184.90

Phrase accuracy

Using command:

./word2vec -train news.2012.en.shuffled-norm1-phrase1 -output vectors-phrase.bin -cbow 1 -size 200 -window 10 -negative 25 -hs 0 -sample 1e-5 -threads 12 -binary 1 -iter 15
./compute-accuracy vectors-phrase.bin < questions-phrases.txt
Accuracy(%) Speed(kwords/s)
Neg-25 35.42 139.99
Neg-5 23.45 453.07
HS 12.80 244.34
HS-rebuild 24.88 224.48

Todo

Performance is still worse than NEG mothod. Consider a way to allow multiple codes per word(like this paper, which may improve the performance.