A few questions on training
Author: aakashrkumarCreated Jan 24, 2023Updated Apr 9, 2023
Hi, I've been planning to train this model, I have a tpu pod(v3-128) through trc, which should equate to ~ 5 tb of ram and 2 tb of vram, I had a few questions about how to begin training the model.
- What would be the appropriate dataset to train on? I was currently considering using the pile for pre training, but gathering human feedback for rlhf still seems like a challenge.
- How large of a model would you recommend?
- I saw you mentioned flash attention, are there any drawbacks to using it, because it seems to be practically the best attention
Thanks for all of your implementations, they have been really helpful to learn from
Source: lucidrains/PaLM-rlhf-pytorch