cannot reproduce result?
Performance discrepancy for torch 0.2.0 and 0.4.0
Did adaptive softmax used when running 1B word dataset?
Pre-trained model?
How to run this code on a multi-cpu computer?
暂无评论,来聊聊你的看法吧