Comments (2)
The empirically best number of merge operations likely depends on your dataset and language pair. We have some numbers by @bhaddow on WMT 2018 datasets for EN<->CS,ET,FI, and BLEU differences between 30,000 and 90,000 merge operations are small (although the effect on rare words might be larger), so 30,000 is a good starting point. If your dataset is relatively small (less than 1 million sentence pairs), I'd recommend that you also try out a smaller number of merge operations, since the model is unlikely to learn a useful representation of a subword that only occurs a few times.
If you have related languages, it is preferable to learn BPE units jointly on the source and target language (see https://github.com/rsennrich/subword-nmt#best-practice-advice-for-byte-pair-encoding-in-nmt ). For language pairs such as English-Chinese, there is no need that the number of operations be the same.
from subword-nmt.
Thank you for your prompt, elaborative answer. Yes my data-set is much smaller. So basically it is a value that is decided on empirically isn't it? I'll check with different values. Thank you.
from subword-nmt.
Related Issues (20)
- subword-nmt HOT 3
- learn_bpe.py error HOT 1
- learn_joint_bpe_and_vocab.py for Japanese HOT 1
- BPE-Dropout question HOT 1
- Recover back code file HOT 1
- DeprecationWarning and ResourceWarning: Enable tracemalloc to get the object allocation traceback HOT 2
- No module named apply_bpe HOT 4
- About the vocabulary size HOT 2
- Readme update please HOT 1
- Question about vocabulary filter HOT 2
- Unknown word and vocabulary filter HOT 2
- How to avoid special char like '\t' being split by bpe HOT 3
- applying BPE(Byte Pair Encoding) fails for large Chinese data tokenized with THULAC HOT 1
- Error: invalid line 2 in BPE codes file when running apply_bpe.py HOT 1
- How to decode BPE when applied to machine translation HOT 1
- learn_bpe.py code question HOT 1
- BrokenPipeError: [Errno 32] Broken pipe HOT 3
- Is it possibile to extend a trained BPE model's merge operations? HOT 2
- How to find all valid BPEs for a word? HOT 3
- AttributeError: 'BPE' object has no attribute 'glossaries_regex' HOT 1
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from subword-nmt.