allennlp
tango
Our great sponsors
allennlp | tango | |
---|---|---|
13 | 5 | |
11,337 | 508 | |
- | 5.3% | |
8.4 | 6.5 | |
over 1 year ago | 4 days ago | |
Python | Python | |
Apache License 2.0 | Apache License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
allennlp
-
How to solve ConfigurationError using HuggingFace Token Classifier
No clue. So what I did was google the error. Here's what I found: https://github.com/allenai/allennlp/issues/4319
- AllenNLP will be unmaintained in December
- AllenNLP Is EOL
- Any recommendation for the replacement of the toolkit jiant? [Research] [Discussion]
- Cedille, the largest French language model, open source with a freely accessible playground
-
[P] Cedille, the largest French language model (6b), released in open source
Another aspect we had fun with is dataset filtering. We have run the whole C4 French dataset through the Detoxify classifier to clean it up 🤬
-
Any allennlp users in this sub?
https://github.com/allenai/allennlp/discussions looks active
- Multilingual C4 (mC4) Dataset now released
- C4 dataset released (800GB Common Crawl-derived text; T5 training data)
tango
- AI2 Tango
-
AllenNLP will be unmaintained in December
Maybe we need to re-work the docs if the DAG aspects stick out to you so much. The main functionality is the cache. If you have a complex experiment, you can still write the code as if all the steps were fast, and let them be slow only the first time you run it. The DAG stuff is also nice, but less important.
That said, you could execute sklearn. If that's what your experiment needs, it's the right thing to do. This is why it gives us the flexibility to also support Jax: https://github.com/allenai/tango/pull/313
The DL-specific stuff is in the components we supply. Like the trainer, dataset handling stuff, file formats, and increasingly, https://github.com/allenai/catwalk.
-
AI2 Introduces Tango, A Python Library For Choreographing Machine Learning Research Experiments By Executing A Series Of Steps
Tango ensures you never operate on outdated data by taking care of your intermediate and final outcomes and finding them again when needed.
What are some alternatives?
cedille-ai - ✒️ Cedille is a large French language model (6B), released under an open-source license
spaCy - 💫 Industrial-strength Natural Language Processing (NLP) in Python
fairseq - Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
thinc - 🔮 A refreshing functional take on deep learning, compatible with your favorite libraries
mesh-transformer-jax - Model parallel transformers in JAX and Haiku
catwalk - This project studies the performance and robustness of language models and task-adaptation methods.
lm-evaluation-harness - A framework for few-shot evaluation of language models.
primeqa - The prime repository for state-of-the-art Multilingual Question Answering research and development.
python-sutime - Python wrapper for Stanford CoreNLP's SUTime
haystack - :mag: LLM orchestration framework to build customizable, production-ready LLM applications. Connect components (models, vector DBs, file converters) to pipelines or agents that can interact with your data. With advanced retrieval methods, it's best suited for building RAG, question answering, semantic search or conversational agent chatbots.
PaddleHub - Awesome pre-trained models toolkit based on PaddlePaddle. (400+ models including Image, Text, Audio, Video and Cross-Modal with Easy Inference & Serving)
ai-tools - Simple command-line AI chat assistant built using the OpenAI API