Encoder-only
What is Pre-training? Initializing neural network model parameters through self-supervised learning (unlabeled data).
The goal is to learn a universal language understanding capability.
BERT BERT Pre-training Tasks Masked Language Modeling (MLM) A certain percentage (15%) of tokens in the input sequence are replaced, and the model is tasked with predicting what those original words were.
The 8-1-1 Rule 80% of the selected tokens are replaced with [MASK]: Acts like a cloze test, enabling the model to learn bidirectional semantic context. 10% are replaced with a random token from the vocabulary: Since the model doesn’t know which tokens are random, this improves error-correction capabilities and forces the model to rely on the global context to generate the correct vector representation. 10% remain as the original token: This anchors the model’s representations toward the actual “true” embeddings. Optimization Objective $$\mathcal{L}{MLM} = - \sum{i \in m} \log P(x_i | \tilde{X}; \theta)$$ Where $\tilde{X}$ is the corrupted input sequence, $m$ is the set of chosen masked positions, and $\theta$ represents the model parameters.