Git Submodules in Practice: Add, Clone, Update, Remove, CI Caching and Common Pitfalls

Key takeaways

Git submodules let you embed another Git repository inside yours, pinned to a specific commit. This guide covers add/update/delete workflows, CI setup, common pitfalls (empty folders, detached HEAD), and when monorepos are a better choice.

What Is a Git Submodule?

A Git submodule is a Git repository embedded inside another Git repository. The outer repo (called the superproject) records the exact commit SHA of the inner repo. This gives you:

  • Source code sharing without publishing to a package registry
  • Strict version pinning — you control exactly which commit of the dependency you use
  • Separate commit history and access control for each repo

The cost: operations that feel automatic in a single repo (clone, pull, CI) require extra steps.

superproject/
├── .gitmodules           ← maps submodule paths to their URLs
├── src/
├── vendor/
│   └── shared-lib/       ← this is a gitlink, not a real directory in superproject
└── ...

Adding a Submodule

# Add a submodule at vendor/shared-lib
git submodule add https://github.com/org/shared-lib.git vendor/shared-lib

# This creates/updates .gitmodules and stages the gitlink
git commit -m "chore: add shared-lib submodule"

The resulting .gitmodules file:

[submodule "vendor/shared-lib"]
    path = vendor/shared-lib
    url = https://github.com/org/shared-lib.git
    branch = main   # optional: track a branch

The superproject now contains a gitlink — an entry in the tree that points to a specific SHA in the child repo. It is not a regular file or directory from Git’s perspective.

This single fact explains most submodule behavior. The superproject never stores the submodule’s files, only “this path should be at commit a1b2c3d of that repository”. Everything else is a consequence: a fresh clone has an empty directory until the submodule is fetched, a git pull in the parent moves the recorded SHA but does not move the files in the submodule’s working tree, and git status in the parent reports modified: vendor/shared-lib (new commits) whenever the checked-out commit and the recorded one differ. The branch = main line does not change what is pinned; it only tells git submodule update --remote which branch to fetch when you ask to move the pin forward.

git submodule add also clones the child repository into .git/modules/vendor/shared-lib and leaves a small .git file in the submodule directory pointing there. Prefer an HTTPS or SSH URL that works for everyone who clones; a relative URL such as ../shared-lib.git resolves against the superproject’s own remote, which is convenient when repositories are mirrored or accessed both over SSH and HTTPS.


Cloning a Repo with Submodules

Regular git clone does not fetch submodules:

# All-in-one: clone and initialize all submodules
git clone --recurse-submodules https://github.com/org/main-app.git

# If you already cloned without --recurse-submodules:
git submodule update --init --recursive

After git pull that updates a submodule pointer:

git pull
git submodule update --init --recursive   # update to the new SHA

Put this in your README. Every developer who doesn’t know about submodules will hit the “empty folder” problem exactly once, then need this command.

Forgetting the update after a pull is the more dangerous variant, because nothing looks empty. The submodule still contains the old commit, the build uses old code, and git status shows the submodule as modified. If someone then runs git commit -a or git add ., they commit the old SHA back and silently revert a teammate’s bump. git config --global submodule.recurse true makes pull, checkout and switch update submodules automatically and removes most of this class of mistake; git config --global status.submoduleSummary true makes git status show which commits differ.


Updating a Submodule to Latest

The superproject does not automatically follow the latest commit in the child repo. You bump it deliberately:

# Enter the submodule
cd vendor/shared-lib

# Get the latest commits
git fetch origin
git checkout main
git pull

# Go back and record the new SHA
cd ../..
git add vendor/shared-lib
git status
# modified: vendor/shared-lib (new commits)
git commit -m "chore: bump shared-lib to latest main"
git push

The git add vendor/shared-lib step does not add files; it records the submodule’s current HEAD commit as the new gitlink SHA. Before committing, git diff --cached --submodule=log lists the child commits you are pulling in, which is worth including in the commit message or pull request, since a one-line SHA change can hide a large upstream change. Treat a bump like a dependency upgrade: review it and run the tests.

Your teammates then:

git pull
git submodule update --init --recursive

Update All Submodules at Once

# Fetch and merge latest for all submodules tracking a branch
git submodule update --remote --merge

# Then review, add, and commit
git add .
git commit -m "chore: bump all submodules"

--remote fetches each submodule and moves it to the tip of its configured branch (branch in .gitmodules, otherwise the remote’s default branch), and --merge merges into any local work instead of detaching; --rebase is the alternative. git add . is used here for brevity, but it stages everything else in the working tree too. Adding the submodule paths explicitly, or reviewing git status first, avoids committing unrelated changes along with the bump.


Removing a Submodule

Git has no single git submodule remove command. The steps are:

# 1. Remove from .gitmodules
git config -f .gitmodules --remove-section submodule.vendor/shared-lib

# 2. Stage the .gitmodules change
git add .gitmodules

# 3. Remove the submodule from the index and working tree
git rm -f vendor/shared-lib

# 4. Clean up the .git/modules cache
rm -rf .git/modules/vendor/shared-lib

# 5. Commit
git commit -m "chore: remove shared-lib submodule"

These manual steps come from older Git versions. Since Git 1.8.5, git rm understands submodules and removes the .gitmodules entry for you, so on any current Git the shorter sequence is:

git submodule deinit -f vendor/shared-lib   # clear the working tree and .git/config entry
git rm vendor/shared-lib                    # remove the gitlink and the .gitmodules section
rm -rf .git/modules/vendor/shared-lib       # optional: drop the cached clone
git commit -m "chore: remove shared-lib submodule"

deinit removes the submodule’s entry from .git/config and empties its directory, which matters for teammates too: after pulling the removal, their .git/config may still mention the submodule until they run git submodule sync or clean it up, which is harmless but can confuse later re-adds.


CI Configuration

Every CI run needs to initialize submodules. Two approaches:

GitHub Actions

# .github/workflows/ci.yml
jobs:
  build:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          submodules: 'recursive'    # initializes all submodules
          token: ${{ secrets.GH_PAT }}  # required for private submodules

The default GITHUB_TOKEN that actions/checkout uses is scoped to the repository running the workflow, so it can fetch public submodules but not private ones in other repositories; that is why a separate token or SSH key is needed. The failure message is usually a generic fatal: could not read Username for 'https://github.com' or Repository not found, which does not mention submodules at all. Submodule URLs using SSH ([email protected]:org/lib.git) need the ssh-key input instead of token.

Shell Script

# If you use a manual checkout step
git checkout $COMMIT_SHA
git submodule update --init --recursive --depth 1  # --depth 1 for faster CI

# Or in a Makefile:
setup:
	git submodule update --init --recursive

Caching Submodules in CI

- name: Cache submodule data
  uses: actions/cache@v4
  with:
    path: vendor/
    key: submodules-${{ hashFiles('.gitmodules') }}-${{ github.sha }}
    restore-keys: |
      submodules-${{ hashFiles('.gitmodules') }}-

Include the submodule directory in the cache key based on .gitmodules — this invalidates the cache when the submodule URL or path changes.

Be careful with what this cache actually buys. Because the key contains github.sha, it never matches exactly on a new commit; it restores the most recent entry through restore-keys, and then the checkout or submodule update step still has to fetch whatever changed. .gitmodules also does not change when a submodule is bumped, since the pinned SHA lives in the tree, not in that file, so the key does not track versions. Caching pays off mainly for large submodules on slow networks; for typical ones, a shallow fetch (--depth 1, or fetch-depth: 1 with actions/checkout) is simpler and often just as fast. One more shallow-clone detail: fetching a pinned SHA with --depth 1 requires the server to allow fetching arbitrary reachable commits, which GitHub and GitLab do, but some self-hosted servers do not, failing with Fetched in submodule path ..., but it did not contain <sha>.


Common Pitfalls

Detached HEAD in the Submodule

When you enter a submodule after git submodule update, you’re in detached HEAD — you’re not on any branch:

cd vendor/shared-lib
git status
# HEAD detached at a1b2c3d

If you make changes here without creating a branch first, those commits may be lost when you update again:

# Before making changes in a submodule:
git checkout -b my-feature   # create a branch
# ... make changes, commit ...
git push origin my-feature   # push to the child repo's remote

Detached HEAD is the intended state, not a malfunction: the superproject asks for a specific commit, and Git checks out exactly that commit without any branch. Commits made there are not lost immediately, but nothing refers to them, and the next git submodule update moves HEAD back to the recorded SHA, leaving them reachable only through git reflog until garbage collection removes them. The workflow that avoids surprises is: create a branch inside the submodule, commit, push the branch to the child repository, open and merge a pull request there, and only then bump the pin in the superproject to a commit that exists on the child’s remote. Skipping the push is the cause of the CI failure described in the FAQ below.

I find that most teams who adopt submodules have one person who understands them and several who work around them. When that happens, it is usually a sign to reconsider: submodules work best when the child repository changes rarely and the bump is a deliberate event, not when developers edit both repositories every day.

Permission Errors in CI

Private submodules require the CI runner to have access to the child repo:

# Option 1: SSH deploy key on the child repo (read-only)
# Add the public key to the child repo's Deploy Keys

# Option 2: GitHub token with repo scope
# Use secrets.GH_PAT in the checkout action

# Option 3: Machine user with access to all repos
# Create a bot account, add it as collaborator on child repos

The .git/modules Cache

After git rm -f vendor/shared-lib, the .git/modules/vendor/shared-lib directory still exists. If you re-add the same submodule path later (even at a different URL), Git uses the cached data and may fail. Always clean it:

rm -rf .git/modules/vendor/shared-lib

Comparing Multi-Repo Strategies

ApproachBest whenDownsides
Git submoduleStrict version pinning, private source-level sharing, no registryMore complex clone/CI, detached HEAD confusion
Git subtreeWant to merge history into one repo, rare upstream syncsHeavier upstream sync workflow
npm/pip/cargo packagePublished library, semantic versioning, public or private registryNeeds publish step for every change
MonorepoTeams change multiple packages together, unified CI pipelineRepo size grows, permission design more complex

When submodules make sense:

  • Shared protobuf/schema files used by services in different languages
  • Internal C++ library shared between firmware and desktop apps
  • Test data or fixtures that are maintained as their own project

When to prefer packages: if you’re sharing JavaScript/TypeScript utilities between Node.js services, publish to a private npm registry (GitHub Packages, Verdaccio) instead — standard dependency tools work better than submodules for this use case.

The deciding question is how often the two repositories change together. If most changes to the shared code are made for one consumer and released together with it, a monorepo (or subtree) avoids the two-step commit and the pin-bump ceremony. If the shared code has its own release cycle and several consumers that upgrade independently, a package with semantic versions expresses that better than a raw SHA. Submodules sit in between: exact pinning, no registry, and source available for debugging, which is why they remain common for C and C++ dependencies, schema repositories and firmware projects, where language-level package managers are weaker. For C++ specifically, CMake’s FetchContent or a package manager like vcpkg can pin a Git tag or commit without the submodule workflow at all.


Workflow Checklist

Add to your README.md:

## Setup

Clone with submodules:
```bash
git clone --recurse-submodules https://github.com/org/main-app.git
```

If you already cloned:
```bash
git submodule update --init --recursive
```

After every `git pull`:
```bash
git submodule update --init --recursive
```

Frequently Asked Questions (FAQ)

Q. CI fails to check out the submodule commit my parent repo points to. What went wrong?

A. The parent repo stores only a commit SHA for each submodule. If you committed inside the submodule and pushed the parent, but never pushed the submodule’s own branch, that SHA exists only on your machine and every other clone fails with an error such as fatal: reference is not a tree. Push the child repo first, then the parent; git push --recurse-submodules=check makes Git refuse to push the parent while referenced submodule commits are missing from the child repo’s remote, and =on-demand pushes them for you.