labml.ai on X: "Autocompletion for Python using a Transformer XL model.
Video: https://t.co/DdS331TVxl
This model was trained on open-source Python code and uses simple beam search to make predictions.
This thread describes the how it works.
🧵👇"
Autocompletion for Python using a Transformer XL model.
Video: youtu.be/ZFzxBPBUh0M
This model was trained on open-source Python code and uses simple beam search to make predictions.
This thread describes the how it works.
🧵👇
Autocompletion for Python using a Transformer XL model.
Video: youtu.be/ZFzxBPBUh0M
This model was trained on open-source Python code and uses simple beam search to make predictions.
This thread describes the how it works.
🧵👇
1/
Github: github.com/lab-ml/python_…
This repository contains code for training the model, making predictions with the model, and a simple extension for Visual Studio Code (@code). You can clone this repo and try it out.
2/
We have published the pertained model (80MB) which will get downloaded automatically if you only want to try it. The instructions for loading and running the VSCode extension are in the readme.
5/
The Transformer XL model lets us predict token by token (with cached token embeddings for previous tokens) because of its relative attention mechanism. This improves the performance a lot in the evaluation phase.
7/
We use byte-pair-encoding (BPE) with a vocabulary size of 1,000. Byte pair encoding merges the most common pairs of characters (or tokens) to create new tokens. This sort of evens out the distribution of tokens.
9/
We use a beam search to get suggestions for autocompletion. A greedy prediction will predict token by token, picking the highest probability token at each step, whereas a beam search will maintain the top-k predictions.
10/
This is a bit slow on computers with no GPU. For instance, it takes about 300ms to predict suggestions on my 2018 MacBook Pro. On a 1080 Ti it takes < 100ms.
11/
The sample code used for the video was based on github.com/karpathy/minGP… by @karpathy. This was not present in training data although the model must have seen several multi-head attention models.
-THE END-