T5X 转到 T5X ReadTheDocs 文档页面。T5X 是一个模块化、可组合、研究友好型框架,用于高性能、可配置的…
Go to T5X ReadTheDocs Documentation Page.
T5X is a modular, composable, research-friendly framework for high-performance, configurable, self-service training, evaluation, and inference of sequence models (starting with language) at many scales.
It is essentially a new and improved implementation of the T5 codebase (based on Mesh TensorFlow) in JAX and Flax. To learn more, see the T5X Paper.
Below is a quick start guide for training models with TPUs on Google Cloud. For additional tutorials and background, see the complete documentation.
T5X can be run with XManager on Vertex AI. Vertex AI is a platform for training that creates TPU instances and runs code on the TPUs. Vertex AI will also shut down the TPUs when the jobs terminate. This is signifcantly easier than managing GCE VMs and TPU VM instances.
Follow the pre-requisites and directions to install XManager.
Request TPU quota as required. GCP projects come with 8 cores by default, which is enough to run one training experiment on a single TPU host. If you want to run multi-host training or run multiple trials in parallel, you will need more quota. Navigate to Quotas.
The quota you want is:
Vertex AI APIus-central1Custom model training TPU V2 cores per regionCustom model training TPU V3 cores per regionCustom model training TPU V2 pod cores per regionCustom model training TPU V3 pod cores per region TIP: You won't be able to run single-host experiments with multi-host quota.
(i.e. you can't run tpu_v2=8 using TPU V2 pod)
t5x/scripts/xm_launch.py.As a running example, we use the WMT14 En-De translation which is described in more detail in the Examples section below.
export GOOGLE_CLOUD_BUCKET_NAME=...
export TFDS_DATA_DIR=gs://$GOOGLE_CLOUD_BUCKET_NAME/t5x/data
export MODEL_DIR=gs://$GOOGLE_CLOUD_BUCKET_NAME/t5x/$(date +%Y%m%d)
# Pre-download dataset in multi-host experiments.
tfds build wmt_t2t_translate --data_dir=$TFDS_DATA_DIR
git clone https://github.com/google-research/t5x
cd ./t5x/
python3 ./t5x/scripts/xm_launch.py \
--gin_file=t5x/examples/t5/t5_1_1/examples/base_wmt_from_scratch.gin \
--model_dir=$MODEL_DIR \
--tfds_data_dir=$TFDS_DATA_DIR
Check gs://$GOOGLE_CLOUD_BUCKET_NAME/t5x/ for the output artifacts, which can
be read by TensorBoard.
Note: NVIDIA has released an updated version of this repository with H100 FP8 support and broad GPU performance improvements. Please visit the NVIDIA Rosetta repository for more details and usage instructions.
T5X can be run easily on GPUs either in single-node configurations or multi-node configurations with a SLURM+pyxis cluster. Further instructions at t5x/contrib/gpu. The t5x/contrib/gpu/scripts_gpu folder contains example scripts for pretraining T5X on The Pile and for finetuning on SQuAD and MNLI. These scripts and associated gin configurations also contain additional GPU optimizations for better throughput. More examples and instructions can be found in the NVIDIA Rosetta repository maintained by NVIDIA with H100 FP8 support and broad GPU performance improvements.
Note that all the commands in this document should be run in the commandline of the TPU VM instance unless otherwise stated.
Follow the instructions to set up a Google Cloud Platform (GCP) account and enable the Cloud TPU API.
Note: T5X also works with GPU, please follow instructions in t5x/contrib/gpu if you'd like to use GPU version.
Create a
Cloud TPU VM instance
following
this instruction.
We recommend that you develop your workflow in a single v3-8 TPU (i.e.,
--accelerator-type=v3-8) and scale up to pod slices once the pipeline is
ready. In this README, we focus on using a single v3-8 TPU. See
here to
learn more about TPU architectures.
With Cloud TPU VMs, you ssh directly into the host machine of the TPU VM. You can install packages, run your code run, etc. in the host machine. Once the TPU instance is created, ssh into it with
gcloud alpha compute tpus tpu-vm ssh ${TPU_NAME} --zone=${ZONE}
where TPU_NAME and ZONE are the name and the zone used in step 2.
Install T5X and the dependencies.
git clone --branch=main https://github.com/google-research/t5x
cd t5x
python3 -m pip install -e '.[tpu]' -f \
https://storage.googleapis.com/jax-releases/libtpu_releases.html
Create Google Cloud Storage (GCS) bucket to store the dataset and model checkpoints. To create a GCS bucket, see these instructions.
(optional) If you prefer working with Jupyter/Colab style environment you can setup a custom Colab runtime by following steps from t5x/notebooks.
As a running example, we use the WMT14 En-De translation. The raw dataset is available in TensorFlow Datasets as "wmt_t2t_translate".
T5 casts the translation task such as the following
{'en': 'That is good.', 'de': 'Das ist gut.'}
to the form called "text-to-text":
{'inputs': 'translate English to German: That is good.', 'targets': 'Das ist gut.'}
This formulation allows many different classes of language tasks to be expressed in a uniform manner and a single encoder-decoder architecture can handle them without any task-specific parameters. For more detail, refer to the [T5 paper (Raffel et al. 2019)][t5_paper].
For a scalable data pipeline and an evaluation framework, we use
SeqIO, which was factored out of the [T5
library][t5_github]. A seqio.Task packages together the raw dataset, vocabulary,
preprocessing such as tokenization and evaluation metrics such as
BLEU and provides a
tf.data instance.
[The T5 library][t5_github] provides a number of seqio.Tasks that were used in the
[T5 paper][t5_paper]. In this example, we use wmt_t2t_ende_v003.
Before training or fine-tuning you need to download ["wmt_t2t_translate"] (https://www.tensorflow.org/datasets/catalog/wmt_t2t_translate) dataset first.
# Data dir to save the processed dataset in "gs://data_dir" format.
TFDS_DATA_DIR="..."
# Make sure that dataset package is up-to-date.
python3 -m pip install --upgrade tfds-nightly
# Pre-download dataset.
tfds build wmt_t2t_translate ${TFDS_DATA_DIR}
To run a training job, we use the t5x/train.py script.
# Model dir to save logs, ckpts, etc. in "gs://model_dir" format.
MODEL_DIR="..."
T5X_DIR="..." # directory where the T5X repo is cloned.
TFDS_DATA_DIR="..."
python3 ${T5X_DIR}/t5x/train.py \
--gin_file="t5x/examples/t5/t5_1_1/examples/base_wmt_from_scratch.gin" \
--gin.MODEL_DIR=\"${MODEL_DIR}\" \
--tfds_data_dir=${TFDS_DATA_DIR}
The configuration for this training run is defined in the Gin file base_wmt_from_scratch.gin. Gin-config is a library to handle configurations based on dependency injection. Among many benefits, Gin allows users to pass custom components such as a custom model to the T5X library without having to modify the core library. The custom components section shows how this is done.
While the core library is independent of Gin, it is central to the examples we
provide. Therefore, we provide a short [introduction][gin-primer] to Gin in the
context of T5X. All the configurations are written to a file "config.gin" in
MODEL_DIR. This makes debugging as well as reproducing the experiment much
easier.
In addition to the config.json, model-info.txt file summarizes the model
parameters (shape, names of the axes, partitioning info) as well as the
optimizer states.
To monitor the training in TensorBoard, it is much easier (due to
authentification issues) to launch the TensorBoard on your own machine and not in
the TPU VM. So in the commandline where you ssh'ed into the TPU VM, launch the
TensorBoard with the logdir pointing to the MODEL_DIR.
# NB: run this on your machine not TPU VM!
MODEL_DIR="..." # Copy from the TPU VM.
tensorboard --logdir=${MODEL_DIR}
Or you can launch the TensorBoard inside a Colab. In a Colab cell, run
from google.colab import auth
auth.authenticate_user()
to authorize the Colab to access the GCS bucket and launch the TensorBoard.
%load_ext tensorboard
model_dir = "..." # Copy from the TPU VM.
%tensorboard --logdir=model_dir
We can leverage the benefits of self-supervised pre-training by initializing from one of our pre-trained models. Here we use the T5.1.1 Base checkpoint.
# Model dir to save logs, ckpts, etc. in "gs://model_dir" format.
MODEL_DIR="..."
# Data dir to save the processed dataset in "gs://data_dir" format.
TFDS_DATA_DIR="..."
T5X_DIR="..." # directory where the T5X repo is cloned.
python3 ${T5X_DIR}/t5x/train.py \
--gin_file="t5x/examples/t5/t5_1_1/examples/base_wmt_finetune.gin" \
--gin.MODEL_DIR=\"${MODEL_DIR}\" \
--tfds_data_dir=${TFDS_DATA_DIR}
Note: when supplying a string, dict, list, tuple value, or a bash variable
via a flag, you must put it in quotes. In the case of strings, it requires
escaped quotes (\"<string>\"). For example:
--gin.utils.DatasetConfig.split=\"validation\" or
--gin.MODEL_DIR=\"${MODEL_DIR}\".
Gin makes it easy to change a number of configurations. For example, you can
change the partitioning.PjitPartitioner.num_partitions (overriding
the value in
base_wmt_from_scratch.gin)
to chanage the parallelism strategy and pass it as a commandline arg.
--gin.partitioning.PjitPartitioner.num_partitions=8
To run the offline (i.e. without training) evaluation, you can use t5x/eval.py
script.
EVAL_OUTPUT_DIR="..." # directory to write eval output
T5X_DIR="..." # directory where the t5x is cloned, e.g., ${HOME}"/t5x".
TFDS_DATA_DIR="..."
CHECKPOINT_PATH="..."
python3 ${T5X_DIR}/t5x/eval.py \
--gin_file="t5x/examples/t5/t5_1_1/examples/base_wmt_eval.gin" \
--gin.CHECKPOINT_PATH=\"${CHECKPOINT_PAT
暂无开放 Issues,或尚未同步最近议题。