
I am a software engineer.
Hello 😊,
We learned the core philosophy behind BERT in the introductory chapter of our previous series chapter. In this chapter, we’ll learn more about BERT and another pivotal process in Machine Learning known as General Language Pre-training, GLP.
Let’s start with a little history of General Language Pre-training.
General Language Pre-training (GLP)
The history of General Language Pre-training dates back to 2013. Before then, language and word representations in Natural Language Processing (NLP) tasks were purely based on statistical and rules-based models. For example, rules-based, like if word ends with "ed" → past tense statistical models, like P("the cat sat") = P(the) × P(cat|the) × P(sat|the cat) which tries to calculate the frequency with which each word appears after another in a sentence. Already, we can see how very limited this is, right? 🙂

In the year 2013, which seems like ages now 🥲, a Google researcher (emphasis on Google), Tomas Mikolov, and his team invented the Word2Vec model. This was arguably the first General Language Pre-training (GLP) implementation. The Word2Vec models work by first feeding raw data from sources like Wikipedia, Reddit, etc. Now this data can run into millions of sentences, which is a huge dataset. The model then trains itself to try to understand language by picking a sentence at random and trying to guess a word from the sentence that it’ll mask itself. For example, given the sentence “The car ran on diesel,” Word2Vec might take the word “car” and try to predict nearby words such as “ran” or “diesel” within a fixed context window. The model uses a very simple neural network with a single hidden layer to learn word representations. This is self-supervised learning because the training signal comes directly from the text, without human-labeled data. Over many sentences, the model adjusts its loss to learn vector representations where words with similar contexts have similar meanings.
Then, in 2014, the Computer Science Department at Stanford University released a paper with an improved GLP titled GloVe: Global Vectors for Word Representation. GloVe scans the entire dataset to create a massive table where every word is both a row and a column. Each cell ((X_{ij})) counts how many times word (j) appears near word (i). Instead of a neural network "guessing" words, GloVe uses Weighted Least Squares regression. It tries to make the dot product of two word vectors equal to the logarithm of their co-occurrence count. This simply means that, for example, using the sentence, “The car ran on diesel”, the GloVe model would try to predict the next word after “car” by checking how many times “ran” appeared after car.
Both Word2Vec and GloVe were mostly used for sentiment analysis tasks. Also, for both Word2Vec and GloVe, they both assign a specific vector (number of representations) to each word, which is a core problem. They both fail to distinguish words and context that can have different meanings, like “bank”, which can mean a financial institution or a river bank.

That brings us to 2018, the year of the transformer and the most defining year in modern Artificial Intelligence history. Now in 2018, a team of researchers from the Allen Institute for Artificial Intelligence and the Paul G. Allen School of Computer Science & Engineering, University of Washington came up with an even better approach to NLP and GLP 🤗. They published a paper titled Deep contextualized word representations which proposed Embeddings from Language Models (ELMO). ELMO solved the previous problem with Word2Vec and GloVe about context due to words that have multiple meanings, like “bank”. ELMO works by first being trained on millions of words unsupervised (GLP), just as Word2Vec.
However, it used a bi-directional approach to understand sentences and language. In this way, using the sentence, “We sat by the bank of the river” as an example, to understand what bank means, it first reads it forward and then backwards, in this way reading this example backwards would help the model relate the bank to a river so it understands that in this context bank means river bank rather than a financial institution. Whew 😀, that’s a lot, right? They used Long Short-Term Memory Neural Networks, which are a type of Recurrent Neural Networks we learned about in one of our previous chapters.
ELMO is used for sentiment analysis, Question and Answer tasks (because the model can understand what the user means, their intent, even when they use different words), and Search Engines (So even when you type something not very clear in the search bar, the model can decipher the intent and search appropriately)
So now what is GLP?
General Language Pre-training (GLP or simply Pretraining) is the idea of first training a model on huge amounts of general text, so it learns the structure of language, before adapting it to specific tasks like translation, question-answering, coding, etc.
Pre-training Approaches:
I had to bring us up to speed about the GLP and its history, because in that way, we’ll better understand and appreciate BERT and the incremental growth in Machine Learning.
Before BERT, there were primarily 2 widely-used approaches for pre-training general language representations: an unsupervised feature-based approach and an unsupervised fine-tuning approach.
Unsupervised feature-based approach: This is the approach used by Word2Vec, GLoVe, and ELMO that we just learned. The model starts by training on large unlabelled text, then uses its internal representations as inputs/features, and then trains another model from them as a base model for tasks like classification or sentiment analysis. We can call this approach static, as the internal representations (vectors) cannot be changed after pretraining.
Unsupervised fine-tuning approach: This is the approach proposed earlier that year in 2018 by the GPT research team from OpenAI. Here, the model still gets trained with huge unlabelled text. The fine-tuning approach does something quite new by continuously learning and updating its parameters using Transformers (emphasis on Transformers). So it does not learn once, like the feature-based approach. We can call this approach dynamic, as the internal representations (vectors) can be changed after pretraining. This creates what’s commonly known today as Foundation Models. We’ll talk about this in more detail when we start learning about GPT.
Enter BERT again lol
BERT’s model architecture is a multi-layer bidirectional Transformer encoder. The Transformer architecture is the bedrock of modern LLMs. We learned about Transformers in Chapters 25 and 26. Below is BERT’s architecture.

It starts with Pre-training using Transformers with huge unlabelled texts. And as we learned in our previous chapter, it used the Cloze Task technique by masking a word and then trying to predict it. In that way, it learns grammar, meaning, and relationships between sentences. From the Pre-training step (GLP), we can see that arrows point to different tasks that can be adapted from the Pre-trained model in the Fine-Tuning step. Here we can see 3 tasks:
Multi-Genre Natural Language Inference (MNLI) - Used for sentence-pair classification: entailment, contradiction, or neutral
Named Entity Recognition (NER) - Identifying named entities like people, locations, organizations, dates, etc.
Stanford Question Answering Dataset (SQuAD) → Used for extractive question answering: finding answer spans in text
I guess we have to stop here for this Chapter, it has gotten longer than expected. We’ll continue with BERT in our next chapter, where we’ll understand how it works, and then we’ll build our own BERT model, however small 🤗.
See you in the next one.




