Skip to main content

Command Palette

Search for a command to run...

Code: 20 Million Parameter BERT Model For SQuAD Tasks

AI Series - Chapter 35

Updated
β€’7 min readβ€’View as Markdown
Code: 20 Million Parameter BERT Model For SQuAD Tasks
R

I am a software engineer.

Hello πŸ‘½,

Finally, we'll pretrain and finetune a BERT (Bidirectional Encoder Representations from Transformers) model from scratch in this chapter, brace up 🦾

Before we proceed, I want to believe you've followed our BERT chapters or have prior knowledge of how BERT works, because it would be necessary to understand how to actually code it. We have a few chapters on BERT, but you can brush up on Chapter 34, BERT in its Elements.

Let's get right into πŸƒπŸ»β€β™‚οΈ

Problem

Train a lean minibert model with approximately 20 million parameters and fine-tune it for SQuAD (Q/A) tasks.

Dataset

We’ll use a Tiny Stories dataset from Huggingface with 2 million short stories.

Training dataset: roneneldan/TinyStories

Solution

For this task, we'll use Kaggle because they have easily one of the best free GPU compute environments. I used two Tesla T4 GPUs for this task. You can create a new notebook at https://kaggle.com and select the T4 GPUs as your machine. You can use any other ML text editor, like Jupyter, on your device if it’s powerful enough.

This solution is split into parts: pretraining and fine-tuning sections.

Pretraining

Here is the notebook for the pretraining task: Pretrain notebook

The notebook contains everything, and it's a bit lengthy for me to explain each part in detail, so I'll summarise this into three sections:

1. Data Processing and Tokenization: Here, we first download the tinystories dataset and tokenize it.

After this step, we get 1,999,789 valid sentences.

2. Token Embedding Logic: In this step, we create Masked Language Model (MLM) and Next Sentence Prediction (NSP) classes.

We also create the BERT Embeddings, Multi-head attention, and Transformer classes.

3. Training and Embedding: The classes created in the previous step are used to pretrain the BERT model to create embeddings across the vocabulary size (30522 in our solution). The vocab size here means the unique tokens a model can recognize and generate, like the words in its dictionary. Embedding here gives meaning to each token (token here can be a word, subword, etc), represented by a vector.

We run this for 50,000 epochs (iterations) to get our pre-trained BERT model πŸ₯²

After the training, we run a small inference test here to see the performance before any fine-tuning. The results below show the top 5 predictions for the masked word in a sentence, which all look solid, right? We've successfully pretrained a 20 million parameter LLM, miniBERT (L=4, H=256, A=4) πŸ₯‚.

The pretrained miniBERT specs: (L=4, H=256, A=4). This means L=number of multi-attention heads, A=number of attention heads within each multi-attention head, and H=the length of the vector representing each token.

Let's fine-tune this for SQuAD (Q/A) tasks

SQuAD Fine-tuning

Here is the notebook for the fine-tuning task: SQuAD fine-tune notebook

SQuAD (Stanford Question Answering Dataset) is used to fine-tune pretrained models like BERT for question and answering tasks using a passage. A question and a passage are passed, and the model is then expected to predict the span (start and stop position) of the answer from the passage. Right now, our pretrained model can basically try MLM and NSP tasks.

I'll summarize this into 3 steps. You can get the complete notebook above for the complete code.

1. Data Processing and Tokenization: Here, we first download the SQuAD dataset, which consists of about 100 thousand questions and answers, broken into a training set of about 90 thousand and a validation set of about 10 thousand. We then tokenize the dataset.

2. Training: Here, we use the same BERT configuration we used to create the pretrained model. The model learns to predict the start and end positions of an answer span in a passage.

This uses two passes during training: the forward and backward pass.

The forward pass tries to predict the start and end position of the answer in the passage for a question and calculates the loss.

The backward pass, on the other hand, calculates the gradient, which is used to adjust the model parameters based on the loss on the forward pass. For example, if the forward pass predicted a wrong start position, the backward pass would adjust the weights on the model, so next time it can predict the correct span positions for the answer.

3. Evaluation: This step evaluates the model's precision and recall using the validation dataset. The model is simply tested here, with questions from the validation dataset, and the precision and recall are calculated.

Precision basically measures quality, so here it measures the accuracy, like in a predicted answer, how many words were actually part of the correct answer. For example, if the answer is "John Francis," but the model predicted "John Francis is an Architect", while the answer is in the span, it picks words not needed, which would reduce its precision.

Recall, on the other hand, measures quantity, so here it measures frequency, like in a predicted answer, how many correct words did the model return. For example, if the answer is "John Francis", but the model predicted "Francis", it missed the word "John", so its recall score would be 50% less because for the two words, it only correctly predicted one. However, if the model predicted "John Francis is an Architect, the recall score would be a perfect score, because the answer "John Francis" is contained in the span.

Long story short, after evaluation of our 20-million-parameter miniBERT model (L=4, H=256, A=4), fine-tuned with a 100k dataset, our score is Exact Match 12.64%, and the F1 score is about 20%. The exact match is the exact correct answers the model got, while the F1 score measures the precision and recall scores, and calculates the mean of the scores.

Why did our model have a considerably poor score?

Could it be because we did not follow the Golden Ratio as published in the Chinchilla Paper, where the researchers stated that to train an efficient model, for every 1 parameter, it should be trained on 20 tokens? So if you have a 1-million-parameter model, you should train it on at least 20 million tokens; anything less is considered "underschooled or under-trained".

In our case, we have a 20-million-parameter model, and we trained it on at least 200 million tokens from those 2 million stories and still got a poor score. The issue here is simply our model parameter size; a 20-million parameter model can learn just as much πŸ™‚. Keep in mind, the base BERT in the research paper was 110million parameters, so 20 million is a fraction, but just enough for our research purposes. Larger models (large parameters) lead to a strict accuracy improvement, as quoted from the BERT paper.

Whew πŸ₯², I know this is a lot, I actually completed pre-training and fine-tuning of this dataset within a day, with free resources from Kaggle. This is proof that we can do more with more resources. We can be thinking of 100million to 1B parameter models 😎.

This is finally the last series chapter on BERT πŸ₯‚, hope the ride was worth it! Up next, we'll start learning about the holy grail of LLMs, Generative Pretrained Transformers (GPT), gear up πŸ‘½

⬅️ Previous Chapter: The Distillation Dilemma: Anthropic, OpenAI, and the Security Predicament for Frontier AI Labs

➑️ Next Chapter: BERT, Dense Passage Retrieval, and how RAG was invented