Back to BlogMLOps

DagsHub — An All-in-One MLOps Platform

💡

DagsHub = GitHub + DVC + MLflow + Label Studio

An all-in-one MLOps platform for managing code, data, experiments, models, and collaboration.


Overview

What you will learn

  • Introduction to DagsHub
  • Environment setup & Git connection
  • DVC (Data Version Control): versioning large datasets + pipelines
  • MLflow on DagsHub: experiment tracking + model registry
  • Demo: EN → VI machine translation with a Transformer (HuggingFace)
  • Best practices & advanced features
  • Troubleshooting

Website


Part 1 — Introduction to DagsHub

1.1 What is DagsHub?

DagsHub is a comprehensive MLOps platform for Data Scientists and ML Engineers. Built as “GitHub for Machine Learning,” DagsHub integrates everything you need to manage the full AI/ML project lifecycle on a single platform.

1.2 Why do you need DagsHub?

Common ML project problems that DagsHub solves:

  • Data versioning: You cannot track changes to large datasets with ordinary Git
  • Experiment chaos: You run many experiments but struggle to find the best config
  • Model management: There is no organized place to store/deploy models
  • Collaboration: It is hard to work as a team with data/model files
  • Reproducibility: It is hard to reproduce results

1.3 Key features

FeatureDescription
Git RepositoryStore code like GitHub, with support for branches, PRs, and issues
Data Versioning (DVC)Track large datasets; store them on S3/GCS/Azure/DagsHub Storage
Experiment TrackingMLflow-compatible; track metrics, params, and artifacts
Model RegistryManage model versions and stages (Staging/Production)
Data AnnotationIntegrates Label Studio for data annotation
Notebook DiffingVisually compare Jupyter notebooks
CI/CD/CTIntegrates with GitHub Actions and GitLab CI
CollaborationShare repos, datasets, and experiments with your team

1.4 Comparison with other platforms

PlatformDagsHub vs ...
GitHubDagsHub adds data versioning, experiment tracking, and ML-specific features
MLflow (standalone)DagsHub hosts an MLflow server + adds Git/DVC integration
Weights & BiasesDagsHub is open-source friendly, self-hostable, and integrates DVC
Neptune.aiDagsHub has a broader free tier and better Git integration
DVC StudioDagsHub = DVC Studio + extra features (annotation, model registry)

Part 2 — Installation & setup

2.1 Create a DagsHub account

  1. Go to https://dagshub.com/user/sign_up
  2. Sign up with email or connect a GitHub/GitLab account
  3. Confirm your email and complete your profile
  4. You can connect GitHub to import existing repositories

2.2 Set up the Python environment

Requirements: Python >= 3.8, pip or conda

# Create a virtual environment
python -m venv dagshub_env
source dagshub_env/bin/activate # Linux/Mac
dagshub_env\Scripts\activate # Windows

# Install the required libraries
pip install dagshub dvc mlflow torch transformers
pip install sacrebleu sentencepiece datasets

2.3 Connect DagsHub with Git

# Install Git (if not already installed)
git config --global user.name "Your Name"
git config --global user.email "your@email.com"

# Clone the repo from DagsHub
git clone https://dagshub.com/<username>/<repo-name>.git
cd <repo-name>

2.4 Get an Access Token

  1. Go to Settings > Security > Access Tokens on DagsHub
  2. Click "Generate new token"
  3. Name the token and choose permissions (read/write)
  4. Copy the token and store it securely (it will not be shown again)
# Set credentials in the terminal
export DAGSHUB_USER_TOKEN=<your_token> # Linux/Mac
set DAGSHUB_USER_TOKEN=<your_token> # Windows CMD

🔒

Security note: Do not commit tokens to a Git repository.

Use a .env file or environment variables, and add .env to .gitignore immediately.


Part 3 — Data Version Control (DVC)

3.1 What is DVC and why use it?

DVC (Data Version Control) is an open-source tool for version control of large data and models. It works alongside Git: metadata lives in Git, while the actual data/models live in remote storage.

3.2 Initialize DVC in a project

# Initialize DVC (in a folder that already has git init)
dvc init

# Configure DagsHub as DVC remote storage
dvc remote add origin https://dagshub.com/<username>/<repo>.dvc
dvc remote modify origin --local auth basic
dvc remote modify origin --local user <dagshub_username>
dvc remote modify origin --local password <access_token>

3.3 Track data files with DVC

# Add the dataset to DVC tracking
dvc add data/raw/train.csv
dvc add data/raw/test.csv

# Commit the metadata to Git
git add data/raw/train.csv.dvc data/raw/test.csv.dvc .gitignore
git commit -m "Add training data via DVC"

# Push data to DagsHub storage
dvc push

3.4 Pipelines with DVC

# Create DVC pipeline stages in dvc.yaml
dvc run -n preprocess \
 -d src/preprocess.py -d data/raw/train.csv \
 -o data/processed/train_processed.csv \
 python src/preprocess.py

# Run the full pipeline
dvc repro

# View the pipeline DAG
dvc dag

3.5 Pull data from DagsHub

# Pull all data
dvc pull

# Pull data for a specific version
git checkout <commit-hash>
dvc checkout

Part 4 — Experiment Tracking with MLflow

4.1 Connect MLflow to DagsHub

DagsHub hosts an MLflow server for every repository. You do not need to install a separate server.

import dagshub
import mlflow

# Method 1: Auto-configure with dagshub.init()
dagshub.init(repo_owner="<username>", repo_name="<repo>", mlflow=True)

# Method 2: Set environment variables manually
import os
os.environ["MLFLOW_TRACKING_URI"] = "https://dagshub.com/<username>/<repo>.mlflow"
os.environ["MLFLOW_TRACKING_USERNAME"] = "<username>"
os.environ["MLFLOW_TRACKING_PASSWORD"] = "<token>"

4.2 Log basic experiments

with mlflow.start_run(run_name="baseline_model"):
	# Log parameters
	mlflow.log_param("learning_rate", 0.001)
	mlflow.log_param("batch_size", 32)
	mlflow.log_param("num_epochs", 10)

	# Training loop ...
	for epoch in range(num_epochs):
		train_loss = train_one_epoch(model, data)
		val_bleu = evaluate(model, val_data)

		# Log metrics every epoch
		mlflow.log_metric("train_loss", train_loss, step=epoch)
		mlflow.log_metric("val_bleu", val_bleu, step=epoch)

	# Log model artifact
	mlflow.pytorch.log_model(model, "translation_model")

4.3 Compare experiments in the UI

  1. Open the repository on DagsHub → "Experiments" tab
  2. Select multiple runs to compare
  3. View metric charts over time
  4. Filter runs by parameters or metrics
  5. Download a comparison CSV for offline analysis

Part 5 — Demo: Machine translation with a Transformer

About the demo task

  • Task: Translate text EN → VI
  • Dataset: OPUS / WMT or a custom dataset
  • Model: Transformer (Helsinki-NLP/opus-mt-en-vi from HuggingFace)
  • Stack: PyTorch + Transformers + DagsHub + MLflow + DVC

5.1 Project structure

machine_translation/
├── .dvc/ # DVC config
├── data/
│ ├── raw/ # Raw data (tracked by DVC)
│ │ ├── train.csv
│ │ └── test.csv
│ └── processed/ # Processed data
├── src/
│ ├── preprocess.py # Data preprocessing
│ ├── train.py # Training script
│ └── evaluate.py # Evaluation script
├── models/ # Saved models (tracked by DVC)
├── notebooks/ # Jupyter notebooks
├── dvc.yaml # DVC pipeline
├── params.yaml # Hyperparameters
├── requirements.txt
└── README.md

5.2 Step 1: Initialize the project and DagsHub

# 1. Create a new repo on the DagsHub UI
# 2. Clone it locally
git clone https://dagshub.com/<username>/machine-translation.git
cd machine-translation

# 3. Initialize DVC
dvc init
git add .dvc .gitignore
git commit -m "Initialize DVC"

# 4. Set up the DVC remote
dvc remote add -d origin https://dagshub.com/<username>/machine-translation.dvc
dvc remote modify origin --local auth basic
dvc remote modify origin --local user <username>
dvc remote modify origin --local password <token>

5.3 Step 2: Prepare the data (preprocess.py)

# src/preprocess.py
from datasets import load_dataset
import pandas as pd
import os

def prepare_data():
	# Load the EN-VI dataset from HuggingFace
	dataset = load_dataset("Helsinki-NLP/opus-100", "en-vi")

	# Convert to a DataFrame
	train_data = []
	for item in dataset["train"]:
		train_data.append({
			"en": item["translation"]["en"],
			"vi": item["translation"]["vi"]
		})

	df = pd.DataFrame(train_data[:50000]) # Take 50k samples
	df.to_csv("data/raw/train.csv", index=False)
	print(f"Saved {len(df)} training samples")

if __name__ == "__main__":
	prepare_data()

# Run preprocessing
# python src/preprocess.py
# Track data with DVC
# dvc add data/raw/train.csv
# git add data/raw/train.csv.dvc
# git commit -m "Add training dataset"
# dvc push

5.4 Step 3: Training script (train.py)

# src/train.py
import dagshub
import mlflow
import torch
from transformers import MarianMTModel, MarianTokenizer
from torch.utils.data import DataLoader, Dataset
import pandas as pd
import yaml

# Initialize DagsHub tracking
dagshub.init(
	repo_owner="<username>",
	repo_name="machine-translation",
	mlflow=True
)

# Load hyperparameters from params.yaml
with open("params.yaml") as f:
	params = yaml.safe_load(f)

class TranslationDataset(Dataset):
	def __init__(self, df, tokenizer, max_len=128):
		self.src = df["en"].tolist()
		self.tgt = df["vi"].tolist()
		self.tokenizer = tokenizer
		self.max_len = max_len

	def __len__(self):
		return len(self.src)

	def __getitem__(self, idx):
		encoding = self.tokenizer(
			self.src[idx],
			text_target=self.tgt[idx],
			max_length=self.max_len,
			truncation=True,
			padding="max_length",
			return_tensors="pt"
		)
		return {k: v.squeeze() for k, v in encoding.items()}

def train():
	model_name = "Helsinki-NLP/opus-mt-en-vi"
	tokenizer = MarianTokenizer.from_pretrained(model_name)
	model = MarianMTModel.from_pretrained(model_name)

	df = pd.read_csv("data/raw/train.csv")
	dataset = TranslationDataset(df, tokenizer)
	loader = DataLoader(
		dataset,
		batch_size=params["batch_size"],
		shuffle=True
	)

	optimizer = torch.optim.AdamW(
		model.parameters(),
		lr=params["learning_rate"]
	)

	with mlflow.start_run(run_name=params["run_name"]):
		mlflow.log_params(params)

		for epoch in range(params["num_epochs"]):
			model.train()
			total_loss = 0

			for batch in loader:
				outputs = model(**batch)
				loss = outputs.loss
				loss.backward()

				optimizer.step()
				optimizer.zero_grad()
				total_loss += loss.item()

			avg_loss = total_loss / len(loader)
			mlflow.log_metric("train_loss", avg_loss, step=epoch)
			print(f"Epoch {epoch+1}: Loss = {avg_loss:.4f}")

		# Save and log the model
		model.save_pretrained("models/final")
		tokenizer.save_pretrained("models/final")
		mlflow.log_artifacts("models/final", "model")

if __name__ == "__main__":
	train()

5.5 Step 4: params.yaml

# params.yaml — Hyperparameters
run_name: "transformer_en_vi_v1"
learning_rate: 2e-5
batch_size: 16
num_epochs: 3
max_length: 128
warmup_steps: 500
model_name: "Helsinki-NLP/opus-mt-en-vi"

5.6 Step 5: Evaluation script (evaluate.py)

# src/evaluate.py
import mlflow
from transformers import pipeline
from sacrebleu.metrics import BLEU
import pandas as pd

def evaluate_model(model_path, test_file):
	translator = pipeline(
		"translation",
		model=model_path,
		device=-1
	) # CPU

	df = pd.read_csv(test_file)
	bleu = BLEU()

	predictions = []
	references = []

	for _, row in df.head(500).iterrows():
		pred = translator(row["en"])[0]["translation_text"]
		predictions.append(pred)
		references.append([row["vi"]])

	score = bleu.corpus_score(predictions, references)
	print(f"BLEU Score: {score.score:.2f}")

	with mlflow.active_run():
		mlflow.log_metric("bleu_score", score.score)

	return score.score

5.7 Step 6: DVC Pipeline (dvc.yaml)

# dvc.yaml
stages:
	preprocess:
		cmd: python src/preprocess.py
		deps:
			- src/preprocess.py
		outs:
			- data/raw/train.csv
	train:
		cmd: python src/train.py
		deps:
			- src/train.py
			- data/raw/train.csv
			- params.yaml
		params:
			- learning_rate
			- batch_size
			- num_epochs
		outs:
			- models/final
		metrics:
			- metrics/train_metrics.json
	evaluate:
		cmd: python src/evaluate.py
		deps:
			- src/evaluate.py
			- models/final
			- data/raw/test.csv
		metrics:
			- metrics/eval_metrics.json:
					cache: false

# Run the full pipeline
# dvc repro

5.8 Step 7: Run experiments with different configs

# Try a different learning rate
dvc exp run --set-param learning_rate=5e-5

# Try a different batch size
dvc exp run --set-param batch_size=32

# Compare experiments
dvc exp show

# Push experiments to DagsHub
dvc exp push origin

Part 6 — Model Registry

6.1 Register a model after training

import mlflow

# Register the model in the registry
run_id = mlflow.last_active_run().info.run_id
model_uri = f"runs:/{run_id}/model"

mlflow.register_model(
	model_uri=model_uri,
	name="en-vi-transformer"
)

6.2 Manage model stages

from mlflow.tracking import MlflowClient

client = MlflowClient()

# Transition the model to Staging
client.transition_model_version_stage(
	name="en-vi-transformer",
	version=1,
	stage="Staging"
)

# After tests pass, promote to Production
client.transition_model_version_stage(
	name="en-vi-transformer",
	version=1,
	stage="Production"
)

6.3 Load a model from the Registry for inference

# Load the model from the Production stage
model = mlflow.pyfunc.load_model(
	model_uri="models:/en-vi-transformer/Production"
)

# Inference
result = model.predict(["Hello, how are you?"])
print(result) # [("Xin chào, bạn có khỏe không?")]

Part 7 — Collaboration & Best Practices

7.1 Teamwork with DagsHub

  • Create a private repository and invite collaborators via Settings > Members
  • Use branches for each experiment or feature
  • Create Pull Requests to review code and experiments before merging
  • Use Issues to track bugs, todos, and discussions
  • Compare experiments in a PR so improvements are clearly visible

7.2 .gitignore and .dvcignore

# .gitignore — things that should NOT be pushed to Git
__pycache__/
- .pyc
.env
models/
data/raw/
data/processed/
- .log

# .dvcignore — ignored by DVC
.git
- .pyc
__pycache__

7.3 CI/CD with GitHub Actions

# .github/workflows/train.yml
name: Train and Evaluate

on:
	push:
		branches: [main]
		paths: [src/**, params.yaml]

jobs:
	train:
		runs-on: ubuntu-latest
		steps:
			- uses: actions/checkout@v3

			- name: Setup Python
				uses: actions/setup-python@v4
				with: { python-version: "3.10" }

			- name: Install dependencies
				run: pip install -r requirements.txt

			- name: Pull data from DagsHub
				env:
					DAGSHUB_TOKEN: $ secrets.DAGSHUB_TOKEN 
				run: dvc pull

			- name: Run pipeline
				run: dvc repro

7.4 Best Practices

CategoryBest Practice
NamingGive runs meaningful names: "transformer_lr2e5_bs16_epoch3"
ParamsAlways use params.yaml instead of hardcoding hyperparameters
TaggingTag experiments with version, dataset, and architecture
MetricsLog metrics every epoch, not only at the end
ArtifactsLog the model, tokenizer, and confusion matrix
ReproducibilitySet a random seed and log it as a param
DocumentationWrite a full description for every experiment run
Data lineageAlways link a model to the dataset version used for training

Part 8 — Advanced features

8.1 DagsHub Storage API (S3-compatible)

import boto3

s3 = boto3.client(
	"s3",
	endpoint_url="https://dagshub.com/s3",
	aws_access_key_id="<username>",
	aws_secret_access_key="<token>"
)

# List files
response = s3.list_objects(Bucket="<username>/<repo>")

# Upload file
s3.upload_file(
	"local_file.csv",
	"<username>/<repo>",
	"path/in/repo.csv"
)

8.2 DagsHub Dataset Streaming

import dagshub.data_engine as de

# Connect to the dataset
ds = de.init("<username>/<repo>", "data/raw")

# Query and filter
filtered = ds.filter(lambda x: x.file_size > 1000)\
	.select(["filename", "label", "split"])

# Download only the files you need
filtered.download()

8.3 Hyperparameter Search with DVC Experiments

# Run a grid search
dvc exp run --queue \
 --set-param learning_rate=1e-5 \
 --set-param batch_size=16

dvc exp run --queue \
 --set-param learning_rate=2e-5 \
 --set-param batch_size=32

dvc exp run --queue \
 --set-param learning_rate=5e-5 \
 --set-param batch_size=64

# Run all queued experiments (in parallel)
dvc queue start --jobs 3

8.4 Custom MLflow Callbacks for HuggingFace

from transformers import TrainerCallback
import mlflow

class DagsHubCallback(TrainerCallback):
	def on_log(self, args, state, control, logs=None, **kwargs):
		if state.is_local_process_zero and logs:
			for k, v in logs.items():
				if isinstance(v, (int, float)):
					mlflow.log_metric(k, v, step=state.global_step)

# Use with the HuggingFace Trainer
trainer = Seq2SeqTrainer(
	model=model,
	args=training_args,
	train_dataset=train_dataset,
	eval_dataset=eval_dataset,
	callbacks=[DagsHubCallback()]
)

8.5 Notebook Versioning and Diffing

  • Push Jupyter notebooks to DagsHub like ordinary code
  • DagsHub shows a clean notebook diff (cell-by-cell)
  • View output changes, metric changes, and code changes separately
  • Especially useful when reviewing notebooks in Pull Requests

Part 9 — Tips, Tricks & Troubleshooting

9.1 Common errors and how to fix them

ErrorSolution
dvc push: 403 ForbiddenRecheck the token and make sure it has write permission
MLflow: Cannot connectCheck MLFLOW_TRACKING_URI and run dagshub.init() first
dvc pull: No dataRun git pull first so you have the latest .dvc files
Large model push slowUse dvc push --jobs 4 for parallel upload
Experiment not showingMake sure mlflow.end_run() is called, or use a context manager

9.2 Useful tips

  • Use tags in MLflow to group experiments: mlflow.set_tag("model_type", "transformer")
  • Log system metrics automatically: mlflow.autolog()
  • Use dvc metrics show to view metrics directly in the terminal
  • Set a default remote in .dvc/config so you do not have to specify it every time
  • Create a DagsHub Organization to manage team repos more effectively
  • Use DagsHub Webhooks to trigger actions when an experiment finishes

9.3 An optimal workflow for an ML project

  1. Create a repo on DagsHub → clone → dvc init
  2. Prep data → dvc add → dvc push → git commit
  3. Write code → dagshub.init() in the training script
  4. Run experiments with dvc exp run → compare them in the UI
  5. Register the best model → promote it to Production
  6. Open a PR → review experiment results → merge

9.4 Further learning resources

Huỳnh Phước Nguyên

Written by Huỳnh Phước Nguyên

AI Engineer, BK Hightech

Ready to build something great?

Tell us about your project and we'll get back to you within a day.

Get in Touch