dvc
aim
Our great sponsors
dvc | aim | |
---|---|---|
108 | 70 | |
13,032 | 4,711 | |
1.6% | 3.9% | |
9.7 | 7.9 | |
about 12 hours ago | about 18 hours ago | |
Python | Python | |
Apache License 2.0 | Apache License 2.0 |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
dvc
-
Why bad scientific code beats code following "best practices"
What you’re describing sounds like DVC (at a higher-ish—80%-solution level).
See pachyderm too.
-
First 15 Open Source Advent projects
10. DVC by Iterative | Github | tutorial
-
Exploring Open-Source Alternatives to Landing AI for Robust MLOps
Platforms such as MLflow monitor the development stages of machine learning models. In parallel, Data Version Control (DVC) brings version control system-like functions to the realm of data sets and models.
- ML Experiments Management with Git
- Ask HN: How do your ML teams version datasets and models?
-
Exploring MLOps Tools and Frameworks: Enhancing Machine Learning Operations
DVC (Data Version Control):
- Evaluate and Track Your LLM Experiments: Introducing TruLens for LLMs
-
[D] Is there a tool to keep track of my ML experiments?
I have been using DVC and MLflow since then DVC had only data tracking and MLflow only model tracking. I can say both are awesome now and maybe the only factor I would like to mention is that IMO, MLflow is a bit harder to learn while DVC is just a git practically.
-
Ask HN: Data Management for AI Training
* User interface for less tech savy people ( e.g just a git like command line is fine for engineers but not for field personell who are not in IT )
I know of tools like https://dvc.org/ but a) they are just layers on top of git b) break appart on huge datasets without a folder hierarchy ( git tree objects just don't work for linear lists of items ) are only useable by IT personell, and require checking out at least a part of the dataset.
Our datasets would be 100.000.000 x 100 MB = 10 PB of raw data. Training data should be delivered to training nodes via network etc.. we just can't have a full checkout of that data...
-
Do you wonder why MLOps is not at the same level as DevOps?
Hey, great find! However, it only explains concepts but not how to actually use any tool. I personally use DVC, but it's more focused on the model development/engineering phase. The different phases of ML are also done independently, which makes it even more difficult for an individual to have exposure to all the different areas. Moreover, the lack of standard tools and best practices makes it difficult, and the fact that every ML problem is different.
aim
-
aim VS cascade - a user suggested alternative
2 projects | 5 Dec 2023
-
Using MLflow(Machine Learning experimentation tracking tool) in Kaggle notebooks with the help of DagsHub
Here is the codebase of aimlflow https://github.com/aimhubio/aimlflow and Aim https://github.com/aimhubio/aim
You can also check out Aim, which has an integration with MLflow, called aimlflow.
-
Effortless image tracking and analysis for 3D segmentation task with Aim
Aim: An easy-to-use & supercharged open-source AI metadata tracker aimstack.io
⭐️ If you find Aim useful, please stop by https://github.com/aimhubio/aim and drop us a star.
-
🦜🔗 Building Multi task AI agent with LangChain and using Aim to trace and visualize the executions
aimstack.io
Here you can find more use-cases: https://github.com/aimhubio/aim/blob/feature/add-langchain-on-docs/docs/source/using/langchain.md
aimstack.io
Hi u/LetGoAndBeReal, Sorry for the inconvenience , here is the link to Aim docs https://github.com/aimhubio/aim/blob/feature/add-langchain-on-docs/docs/source/using/langchain.md feel free to use.
What are some alternatives?
MLflow - Open source platform for the machine learning lifecycle
lakeFS - lakeFS - Data version control for your data lake | Git for data
Activeloop Hub - Data Lake for Deep Learning. Build, manage, query, version, & visualize datasets. Stream data real-time to PyTorch/TensorFlow. https://activeloop.ai [Moved to: https://github.com/activeloopai/deeplake]
tensorboard - TensorFlow's Visualization Toolkit
delta - An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Trino, and Hive and APIs
ploomber - The fastest ⚡️ way to build data pipelines. Develop iteratively, deploy anywhere. ☁️
git-submodules - Git Submodule alternative with equivalent features, but easier to use and maintain.
guildai - Experiment tracking, ML developer tools
palm-dbt - dbt plugin for Palm CLI
git-lfs - Git extension for versioning large files
metaflow - :rocket: Build and manage real-life ML, AI, and data science projects with ease!