💡
DagsHub = GitHub + DVC + MLflow + Label Studio
An all-in-one MLOps platform for managing code, data, experiments, models, and collaboration.
Overview
What you will learn
- Introduction to DagsHub
- Environment setup & Git connection
- DVC (Data Version Control): versioning large datasets + pipelines
- MLflow on DagsHub: experiment tracking + model registry
- Demo: EN → VI machine translation with a Transformer (HuggingFace)
- Best practices & advanced features
- Troubleshooting
Website
- https://dagshub.com | Used by more than 65,000 data scientists
Part 1 — Introduction to DagsHub
1.1 What is DagsHub?
DagsHub is a comprehensive MLOps platform for Data Scientists and ML Engineers. Built as “GitHub for Machine Learning,” DagsHub integrates everything you need to manage the full AI/ML project lifecycle on a single platform.
1.2 Why do you need DagsHub?
Common ML project problems that DagsHub solves:
- Data versioning: You cannot track changes to large datasets with ordinary Git
- Experiment chaos: You run many experiments but struggle to find the best config
- Model management: There is no organized place to store/deploy models
- Collaboration: It is hard to work as a team with data/model files
- Reproducibility: It is hard to reproduce results
1.3 Key features
| Feature | Description |
|---|---|
| Git Repository | Store code like GitHub, with support for branches, PRs, and issues |
| Data Versioning (DVC) | Track large datasets; store them on S3/GCS/Azure/DagsHub Storage |
| Experiment Tracking | MLflow-compatible; track metrics, params, and artifacts |
| Model Registry | Manage model versions and stages (Staging/Production) |
| Data Annotation | Integrates Label Studio for data annotation |
| Notebook Diffing | Visually compare Jupyter notebooks |
| CI/CD/CT | Integrates with GitHub Actions and GitLab CI |
| Collaboration | Share repos, datasets, and experiments with your team |
1.4 Comparison with other platforms
| Platform | DagsHub vs ... |
|---|---|
| GitHub | DagsHub adds data versioning, experiment tracking, and ML-specific features |
| MLflow (standalone) | DagsHub hosts an MLflow server + adds Git/DVC integration |
| Weights & Biases | DagsHub is open-source friendly, self-hostable, and integrates DVC |
| Neptune.ai | DagsHub has a broader free tier and better Git integration |
| DVC Studio | DagsHub = DVC Studio + extra features (annotation, model registry) |
Part 2 — Installation & setup
2.1 Create a DagsHub account
- Go to https://dagshub.com/user/sign_up
- Sign up with email or connect a GitHub/GitLab account
- Confirm your email and complete your profile
- You can connect GitHub to import existing repositories
2.2 Set up the Python environment
Requirements: Python >= 3.8, pip or conda
# Create a virtual environment
python -m venv dagshub_env
source dagshub_env/bin/activate # Linux/Mac
dagshub_env\Scripts\activate # Windows
# Install the required libraries
pip install dagshub dvc mlflow torch transformers
pip install sacrebleu sentencepiece datasets
2.3 Connect DagsHub with Git
# Install Git (if not already installed)
git config --global user.name "Your Name"
git config --global user.email "your@email.com"
# Clone the repo from DagsHub
git clone https://dagshub.com/<username>/<repo-name>.git
cd <repo-name>
2.4 Get an Access Token
- Go to Settings > Security > Access Tokens on DagsHub
- Click "Generate new token"
- Name the token and choose permissions (read/write)
- Copy the token and store it securely (it will not be shown again)
# Set credentials in the terminal
export DAGSHUB_USER_TOKEN=<your_token> # Linux/Mac
set DAGSHUB_USER_TOKEN=<your_token> # Windows CMD
🔒
Security note: Do not commit tokens to a Git repository.
Use a .env file or environment variables, and add .env to .gitignore immediately.
Part 3 — Data Version Control (DVC)
3.1 What is DVC and why use it?
DVC (Data Version Control) is an open-source tool for version control of large data and models. It works alongside Git: metadata lives in Git, while the actual data/models live in remote storage.
3.2 Initialize DVC in a project
# Initialize DVC (in a folder that already has git init)
dvc init
# Configure DagsHub as DVC remote storage
dvc remote add origin https://dagshub.com/<username>/<repo>.dvc
dvc remote modify origin --local auth basic
dvc remote modify origin --local user <dagshub_username>
dvc remote modify origin --local password <access_token>
3.3 Track data files with DVC
# Add the dataset to DVC tracking
dvc add data/raw/train.csv
dvc add data/raw/test.csv
# Commit the metadata to Git
git add data/raw/train.csv.dvc data/raw/test.csv.dvc .gitignore
git commit -m "Add training data via DVC"
# Push data to DagsHub storage
dvc push
3.4 Pipelines with DVC
# Create DVC pipeline stages in dvc.yaml
dvc run -n preprocess \
-d src/preprocess.py -d data/raw/train.csv \
-o data/processed/train_processed.csv \
python src/preprocess.py
# Run the full pipeline
dvc repro
# View the pipeline DAG
dvc dag
3.5 Pull data from DagsHub
# Pull all data
dvc pull
# Pull data for a specific version
git checkout <commit-hash>
dvc checkout
Part 4 — Experiment Tracking with MLflow
4.1 Connect MLflow to DagsHub
DagsHub hosts an MLflow server for every repository. You do not need to install a separate server.
import dagshub
import mlflow
# Method 1: Auto-configure with dagshub.init()
dagshub.init(repo_owner="<username>", repo_name="<repo>", mlflow=True)
# Method 2: Set environment variables manually
import os
os.environ["MLFLOW_TRACKING_URI"] = "https://dagshub.com/<username>/<repo>.mlflow"
os.environ["MLFLOW_TRACKING_USERNAME"] = "<username>"
os.environ["MLFLOW_TRACKING_PASSWORD"] = "<token>"
4.2 Log basic experiments
with mlflow.start_run(run_name="baseline_model"):
# Log parameters
mlflow.log_param("learning_rate", 0.001)
mlflow.log_param("batch_size", 32)
mlflow.log_param("num_epochs", 10)
# Training loop ...
for epoch in range(num_epochs):
train_loss = train_one_epoch(model, data)
val_bleu = evaluate(model, val_data)
# Log metrics every epoch
mlflow.log_metric("train_loss", train_loss, step=epoch)
mlflow.log_metric("val_bleu", val_bleu, step=epoch)
# Log model artifact
mlflow.pytorch.log_model(model, "translation_model")
4.3 Compare experiments in the UI
- Open the repository on DagsHub → "Experiments" tab
- Select multiple runs to compare
- View metric charts over time
- Filter runs by parameters or metrics
- Download a comparison CSV for offline analysis
Part 5 — Demo: Machine translation with a Transformer
About the demo task
- Task: Translate text EN → VI
- Dataset: OPUS / WMT or a custom dataset
- Model: Transformer (Helsinki-NLP/opus-mt-en-vi from HuggingFace)
- Stack: PyTorch + Transformers + DagsHub + MLflow + DVC
5.1 Project structure
machine_translation/
├── .dvc/ # DVC config
├── data/
│ ├── raw/ # Raw data (tracked by DVC)
│ │ ├── train.csv
│ │ └── test.csv
│ └── processed/ # Processed data
├── src/
│ ├── preprocess.py # Data preprocessing
│ ├── train.py # Training script
│ └── evaluate.py # Evaluation script
├── models/ # Saved models (tracked by DVC)
├── notebooks/ # Jupyter notebooks
├── dvc.yaml # DVC pipeline
├── params.yaml # Hyperparameters
├── requirements.txt
└── README.md
5.2 Step 1: Initialize the project and DagsHub
# 1. Create a new repo on the DagsHub UI
# 2. Clone it locally
git clone https://dagshub.com/<username>/machine-translation.git
cd machine-translation
# 3. Initialize DVC
dvc init
git add .dvc .gitignore
git commit -m "Initialize DVC"
# 4. Set up the DVC remote
dvc remote add -d origin https://dagshub.com/<username>/machine-translation.dvc
dvc remote modify origin --local auth basic
dvc remote modify origin --local user <username>
dvc remote modify origin --local password <token>
5.3 Step 2: Prepare the data (preprocess.py)
# src/preprocess.py
from datasets import load_dataset
import pandas as pd
import os
def prepare_data():
# Load the EN-VI dataset from HuggingFace
dataset = load_dataset("Helsinki-NLP/opus-100", "en-vi")
# Convert to a DataFrame
train_data = []
for item in dataset["train"]:
train_data.append({
"en": item["translation"]["en"],
"vi": item["translation"]["vi"]
})
df = pd.DataFrame(train_data[:50000]) # Take 50k samples
df.to_csv("data/raw/train.csv", index=False)
print(f"Saved {len(df)} training samples")
if __name__ == "__main__":
prepare_data()
# Run preprocessing
# python src/preprocess.py
# Track data with DVC
# dvc add data/raw/train.csv
# git add data/raw/train.csv.dvc
# git commit -m "Add training dataset"
# dvc push
5.4 Step 3: Training script (train.py)
# src/train.py
import dagshub
import mlflow
import torch
from transformers import MarianMTModel, MarianTokenizer
from torch.utils.data import DataLoader, Dataset
import pandas as pd
import yaml
# Initialize DagsHub tracking
dagshub.init(
repo_owner="<username>",
repo_name="machine-translation",
mlflow=True
)
# Load hyperparameters from params.yaml
with open("params.yaml") as f:
params = yaml.safe_load(f)
class TranslationDataset(Dataset):
def __init__(self, df, tokenizer, max_len=128):
self.src = df["en"].tolist()
self.tgt = df["vi"].tolist()
self.tokenizer = tokenizer
self.max_len = max_len
def __len__(self):
return len(self.src)
def __getitem__(self, idx):
encoding = self.tokenizer(
self.src[idx],
text_target=self.tgt[idx],
max_length=self.max_len,
truncation=True,
padding="max_length",
return_tensors="pt"
)
return {k: v.squeeze() for k, v in encoding.items()}
def train():
model_name = "Helsinki-NLP/opus-mt-en-vi"
tokenizer = MarianTokenizer.from_pretrained(model_name)
model = MarianMTModel.from_pretrained(model_name)
df = pd.read_csv("data/raw/train.csv")
dataset = TranslationDataset(df, tokenizer)
loader = DataLoader(
dataset,
batch_size=params["batch_size"],
shuffle=True
)
optimizer = torch.optim.AdamW(
model.parameters(),
lr=params["learning_rate"]
)
with mlflow.start_run(run_name=params["run_name"]):
mlflow.log_params(params)
for epoch in range(params["num_epochs"]):
model.train()
total_loss = 0
for batch in loader:
outputs = model(**batch)
loss = outputs.loss
loss.backward()
optimizer.step()
optimizer.zero_grad()
total_loss += loss.item()
avg_loss = total_loss / len(loader)
mlflow.log_metric("train_loss", avg_loss, step=epoch)
print(f"Epoch {epoch+1}: Loss = {avg_loss:.4f}")
# Save and log the model
model.save_pretrained("models/final")
tokenizer.save_pretrained("models/final")
mlflow.log_artifacts("models/final", "model")
if __name__ == "__main__":
train()
5.5 Step 4: params.yaml
# params.yaml — Hyperparameters
run_name: "transformer_en_vi_v1"
learning_rate: 2e-5
batch_size: 16
num_epochs: 3
max_length: 128
warmup_steps: 500
model_name: "Helsinki-NLP/opus-mt-en-vi"
5.6 Step 5: Evaluation script (evaluate.py)
# src/evaluate.py
import mlflow
from transformers import pipeline
from sacrebleu.metrics import BLEU
import pandas as pd
def evaluate_model(model_path, test_file):
translator = pipeline(
"translation",
model=model_path,
device=-1
) # CPU
df = pd.read_csv(test_file)
bleu = BLEU()
predictions = []
references = []
for _, row in df.head(500).iterrows():
pred = translator(row["en"])[0]["translation_text"]
predictions.append(pred)
references.append([row["vi"]])
score = bleu.corpus_score(predictions, references)
print(f"BLEU Score: {score.score:.2f}")
with mlflow.active_run():
mlflow.log_metric("bleu_score", score.score)
return score.score
5.7 Step 6: DVC Pipeline (dvc.yaml)
# dvc.yaml
stages:
preprocess:
cmd: python src/preprocess.py
deps:
- src/preprocess.py
outs:
- data/raw/train.csv
train:
cmd: python src/train.py
deps:
- src/train.py
- data/raw/train.csv
- params.yaml
params:
- learning_rate
- batch_size
- num_epochs
outs:
- models/final
metrics:
- metrics/train_metrics.json
evaluate:
cmd: python src/evaluate.py
deps:
- src/evaluate.py
- models/final
- data/raw/test.csv
metrics:
- metrics/eval_metrics.json:
cache: false
# Run the full pipeline
# dvc repro
5.8 Step 7: Run experiments with different configs
# Try a different learning rate
dvc exp run --set-param learning_rate=5e-5
# Try a different batch size
dvc exp run --set-param batch_size=32
# Compare experiments
dvc exp show
# Push experiments to DagsHub
dvc exp push origin
Part 6 — Model Registry
6.1 Register a model after training
import mlflow
# Register the model in the registry
run_id = mlflow.last_active_run().info.run_id
model_uri = f"runs:/{run_id}/model"
mlflow.register_model(
model_uri=model_uri,
name="en-vi-transformer"
)
6.2 Manage model stages
from mlflow.tracking import MlflowClient
client = MlflowClient()
# Transition the model to Staging
client.transition_model_version_stage(
name="en-vi-transformer",
version=1,
stage="Staging"
)
# After tests pass, promote to Production
client.transition_model_version_stage(
name="en-vi-transformer",
version=1,
stage="Production"
)
6.3 Load a model from the Registry for inference
# Load the model from the Production stage
model = mlflow.pyfunc.load_model(
model_uri="models:/en-vi-transformer/Production"
)
# Inference
result = model.predict(["Hello, how are you?"])
print(result) # [("Xin chào, bạn có khỏe không?")]
Part 7 — Collaboration & Best Practices
7.1 Teamwork with DagsHub
- Create a private repository and invite collaborators via Settings > Members
- Use branches for each experiment or feature
- Create Pull Requests to review code and experiments before merging
- Use Issues to track bugs, todos, and discussions
- Compare experiments in a PR so improvements are clearly visible
7.2 .gitignore and .dvcignore
# .gitignore — things that should NOT be pushed to Git
__pycache__/
- .pyc
.env
models/
data/raw/
data/processed/
- .log
# .dvcignore — ignored by DVC
.git
- .pyc
__pycache__
7.3 CI/CD with GitHub Actions
# .github/workflows/train.yml
name: Train and Evaluate
on:
push:
branches: [main]
paths: [src/**, params.yaml]
jobs:
train:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- name: Setup Python
uses: actions/setup-python@v4
with: { python-version: "3.10" }
- name: Install dependencies
run: pip install -r requirements.txt
- name: Pull data from DagsHub
env:
DAGSHUB_TOKEN: $ secrets.DAGSHUB_TOKEN
run: dvc pull
- name: Run pipeline
run: dvc repro
7.4 Best Practices
| Category | Best Practice |
|---|---|
| Naming | Give runs meaningful names: "transformer_lr2e5_bs16_epoch3" |
| Params | Always use params.yaml instead of hardcoding hyperparameters |
| Tagging | Tag experiments with version, dataset, and architecture |
| Metrics | Log metrics every epoch, not only at the end |
| Artifacts | Log the model, tokenizer, and confusion matrix |
| Reproducibility | Set a random seed and log it as a param |
| Documentation | Write a full description for every experiment run |
| Data lineage | Always link a model to the dataset version used for training |
Part 8 — Advanced features
8.1 DagsHub Storage API (S3-compatible)
import boto3
s3 = boto3.client(
"s3",
endpoint_url="https://dagshub.com/s3",
aws_access_key_id="<username>",
aws_secret_access_key="<token>"
)
# List files
response = s3.list_objects(Bucket="<username>/<repo>")
# Upload file
s3.upload_file(
"local_file.csv",
"<username>/<repo>",
"path/in/repo.csv"
)
8.2 DagsHub Dataset Streaming
import dagshub.data_engine as de
# Connect to the dataset
ds = de.init("<username>/<repo>", "data/raw")
# Query and filter
filtered = ds.filter(lambda x: x.file_size > 1000)\
.select(["filename", "label", "split"])
# Download only the files you need
filtered.download()
8.3 Hyperparameter Search with DVC Experiments
# Run a grid search
dvc exp run --queue \
--set-param learning_rate=1e-5 \
--set-param batch_size=16
dvc exp run --queue \
--set-param learning_rate=2e-5 \
--set-param batch_size=32
dvc exp run --queue \
--set-param learning_rate=5e-5 \
--set-param batch_size=64
# Run all queued experiments (in parallel)
dvc queue start --jobs 3
8.4 Custom MLflow Callbacks for HuggingFace
from transformers import TrainerCallback
import mlflow
class DagsHubCallback(TrainerCallback):
def on_log(self, args, state, control, logs=None, **kwargs):
if state.is_local_process_zero and logs:
for k, v in logs.items():
if isinstance(v, (int, float)):
mlflow.log_metric(k, v, step=state.global_step)
# Use with the HuggingFace Trainer
trainer = Seq2SeqTrainer(
model=model,
args=training_args,
train_dataset=train_dataset,
eval_dataset=eval_dataset,
callbacks=[DagsHubCallback()]
)
8.5 Notebook Versioning and Diffing
- Push Jupyter notebooks to DagsHub like ordinary code
- DagsHub shows a clean notebook diff (cell-by-cell)
- View output changes, metric changes, and code changes separately
- Especially useful when reviewing notebooks in Pull Requests
Part 9 — Tips, Tricks & Troubleshooting
9.1 Common errors and how to fix them
| Error | Solution |
|---|---|
| dvc push: 403 Forbidden | Recheck the token and make sure it has write permission |
| MLflow: Cannot connect | Check MLFLOW_TRACKING_URI and run dagshub.init() first |
| dvc pull: No data | Run git pull first so you have the latest .dvc files |
| Large model push slow | Use dvc push --jobs 4 for parallel upload |
| Experiment not showing | Make sure mlflow.end_run() is called, or use a context manager |
9.2 Useful tips
- Use tags in MLflow to group experiments:
mlflow.set_tag("model_type", "transformer") - Log system metrics automatically:
mlflow.autolog() - Use
dvc metrics showto view metrics directly in the terminal - Set a default remote in
.dvc/configso you do not have to specify it every time - Create a DagsHub Organization to manage team repos more effectively
- Use DagsHub Webhooks to trigger actions when an experiment finishes
9.3 An optimal workflow for an ML project
- Create a repo on DagsHub → clone →
dvc init - Prep data →
dvc add→dvc push→git commit - Write code →
dagshub.init()in the training script - Run experiments with
dvc exp run→ compare them in the UI - Register the best model → promote it to Production
- Open a PR → review experiment results → merge
9.4 Further learning resources
- Official documentation: https://dagshub.com/docs
- DagsHub Blog: https://dagshub.com/blog — many practical tutorials
- YouTube: the DagsHub channel has video tutorials
- GitHub: https://github.com/DAGsHub — open source tools
- Discord community: Ask questions with the DagsHub team and the community

Written by Huỳnh Phước Nguyên
AI Engineer, BK Hightech
