distributed-llm-pretraining-torchtitan

by Orchestra-Research 12k MIT Updated Jun 16, 2026
distributed-llm-pretraining-torchtitan skill by Orchestra-Research
distributed-llm-pretraining-torchtitan — Workflow skill by Orchestra-Research

distributed-llm-pretraining-torchtitan is an open-source workflow skill for Claude Code and compatible agents, published by Orchestra-Research. Its author describes it as: “Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP). Use when pretraining Llama 3.1, DeepSeek V3, or custom models at scale from 8 to 512+ GPUs with Float8, tor…”. The project has 12k stars on GitHub and is available under the MIT license. Add it to your setup with `/plugin marketplace add Orchestra-Research/AI-Research-SKILLs`.

What distributed-llm-pretraining-torchtitan does

TorchTitan is PyTorch's official platform for large-scale LLM pretraining with composable 4D parallelism (FSDP2, TP, PP, CP), achieving 65%+ speedups over baselines on H100 GPUs.

Installation

Add distributed-llm-pretraining-torchtitan to your agent with:

/plugin marketplace add Orchestra-Research/AI-Research-SKILLs

Always review a skill's source before installing it. This command comes from the skill's public repository; the linked repo is the source of truth for exact setup steps.

What's inside

The SKILL.md for distributed-llm-pretraining-torchtitan is organised into these sections:

  • Quick start
  • Common workflows
  • Workflow 1: Pretrain Llama 3.1 8B on single node
  • Workflow 2: Multi-node training with SLURM
  • Workflow 3: Enable Float8 training for H100s
  • Workflow 4: 4D parallelism for 405B models
  • When to use vs alternatives
  • Common issues
  • Supported models
  • Performance benchmarks (H100)
  • Advanced topics
  • Resources

When to use it

Reach for distributed-llm-pretraining-torchtitan when you want workflow help from your agent without writing the same instructions every session. Load the skill and the agent picks it up automatically for relevant tasks.

Strengths

  • Clear MIT license — safe to read and adapt
  • Ships in Orchestra-Research/AI-Research-SKILLs, an established project with 11,807 GitHub stars
  • Actively maintained (recent commits)

Topics

model architecturedistributed trainingtorchtitanfsdp2tensor parallelpipeline parallelcontext parallelfloat8llamapretrainingaiai-researchclaudeclaude-codeclaude-skillscodexgeminigpt-5grpohuggingface

Frequently asked questions

What does distributed-llm-pretraining-torchtitan do?
Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP). Use when pretraining Llama 3.1, DeepSeek V3, or custom models at scale from 8 to 512+ GPUs with Float8, torch.compile, and distributed checkpointing.
How do I install distributed-llm-pretraining-torchtitan?
Run /plugin marketplace add Orchestra-Research/AI-Research-SKILLs in your agent, then reload your skills. Review the source at https://github.com/Orchestra-Research/AI-Research-SKILLs before installing.
Is distributed-llm-pretraining-torchtitan free to use?
Yes. distributed-llm-pretraining-torchtitan is free and open source under the MIT license, so you can read, run, and adapt it within that license's terms.
Where does distributed-llm-pretraining-torchtitan come from?
distributed-llm-pretraining-torchtitan ships inside Orchestra-Research/AI-Research-SKILLs, a repository that contains 41 catalogued skills in total. The repository's 11,807 GitHub stars apply to that whole collection, not to this skill on its own.

Related skills

More Workflow →

Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention

Trains large language models (2B-462B parameters) using NVIDIA Megatron-Core with advanced parallelism strategies. Use when training models >1B parameters, need maximum GPU efficiency (47% MFU on H100), or require tensor/pipeline/sequence/context/expert parallelism. Production-ready framework used for Nemotron, LLaMA, DeepSeek.

Distributed training orchestration across clusters. Scales PyTorch/TensorFlow/HuggingFace from laptop to 1000s of nodes. Built-in hyperparameter tuning with Ray Tune, fault tolerance, elastic scaling. Use when training massive models across multiple machines or running distributed hyperparameter sweeps.

Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods. Use when fine-tuning large models (7B-70B) with limited GPU memory, when you need to train <1% of parameters with minimal accuracy loss, or for multi-adapter serving. HuggingFace's official library integrated with transformers ecosystem.