Skip to content
Csaba Polyak edited this page Sep 28, 2026 · 1 revision

GitHub Mirror Backup

GitHub Mirror Backup is a lightweight Bash-based backup engine for creating and maintaining local mirrors of GitHub repositories.

It is designed primarily for Linux servers, NAS devices and other always-on systems where GitHub repositories need to exist as an independent local backup.

What does it do?

The script connects to the GitHub API using a Personal Access Token and discovers repositories accessible to the authenticated account.

It then uses Git over SSH to create and maintain local bare mirrors.

                 GitHub
                   │
          ┌────────┴────────┐
          │                 │
       REST API           Git/SSH
          │                 │
   Repository list       Git objects
          │                 │
          └────────┬────────┘
                   │
                   ▼
             Local storage
             ┌───────────┐
             │ repo.git  │
             │ repo.git  │
             │ repo.git  │
             └───────────┘

The API and Git transport are intentionally separated:

  • GitHub API → repository discovery
  • SSH → repository mirroring
  • Local storage → backup destination

The GitHub Personal Access Token is never embedded into Git repository URLs.

Why mirror instead of clone?

A normal clone is primarily intended for working with a repository.

A mirror is different.

git clone --mirror

creates a bare repository containing the repository's Git refs and objects, making it suitable for backup and later restoration.

The mirror can be updated without maintaining a working tree:

git remote update --prune

This makes the backup relatively cheap to maintain after the initial synchronization.

What is backed up?

The Git repository itself is mirrored, including:

  • branches
  • tags
  • Git refs
  • commits
  • trees
  • blobs
  • Git objects

The script also keeps track of repository state such as:

  • public/private visibility
  • forks
  • archived repositories
  • repository ownership
  • repositories that are no longer visible

What is not backed up?

GitHub contains considerably more data than the Git repository itself.

The current version does not back up:

  • Issues
  • Pull Requests and their discussions
  • GitHub Actions artifacts
  • Releases metadata
  • GitHub Packages
  • Repository settings
  • Collaborator permissions
  • Webhooks
  • Secrets
  • GitHub Wiki
  • Git LFS objects as a separate backup set

These require additional GitHub API or Git LFS handling and are intentionally outside the scope of the basic mirror engine.

Repository discovery

Repositories are discovered through:

GET /user/repos

The request includes repositories where the authenticated account is:

  • owner
  • collaborator
  • organization member

Pagination is handled automatically, so the script is not limited to the first 100 repositories.

Local repository layout

By default:

backup/
├── project-a.git/
├── project-b.git/
├── project-c.git/
└── _logs/
    ├── github-backup-YYYY-MM-DD_HH-MM-SS.log
    └── ...

Repositories with identical names are stored using their owner name:

alice__website.git
bob__website.git

This prevents one repository from overwriting another.

Safety philosophy

The script follows a few important rules:

Never delete a backup just because GitHub changed.

If a repository disappears from the API results, its local mirror is reported as an orphaned mirror and is kept.

Never overwrite an invalid local repository.

If something exists at the expected path but is not a valid mirror, it is moved aside before a new clone is attempted.

One failed repository should not stop the entire backup.

A failed repository is reported and the remaining repositories continue to be processed.

Typical workflow

First run

Discover repositories
        ↓
Create missing mirrors
        ↓
Report failures
        ↓
Show backup summary

The first run can be relatively expensive because every repository has to be transferred.

Subsequent runs

Discover repositories
        ↓
Check existing mirrors
        ↓
Fetch changes
        ↓
Prune removed refs
        ↓
Report changes

Once the initial mirror exists, subsequent runs normally transfer only changes.

Intended environment

The project is intentionally simple.

It does not require:

  • Docker
  • Python
  • a database
  • GitHub CLI
  • a web interface
  • a proprietary backup service

The core dependencies are standard command-line tools:

bash
git
curl
jq
ssh

This makes it suitable for small servers and NAS systems where installing a large software stack is undesirable.

Next

For installation and first-time setup, see:

[Installation](Installation)

For GitHub authentication:

[GitHub Authentication](GitHub-Authentication)

For scheduled backups:

[Scheduling](Scheduling)

For restoring a repository:

[Recovery](Recovery)

Clone this wiki locally