Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
< Back to tools
H

hercules

> 开发工具
Open source

Gaining advanced insights from Git repository history.

2.8K stars0 likes0 views
WebsiteGitHub

About

Gaining advanced insights from Git repository history.

Hercules

Fast, insightful and highly customizable Git history analysis.

Overview • How To Use • Installation • Contributions • License

-------- Table of Contents ================= * [Overview](#overview) * [Installation](#installation) * [Build from source](#build-from-source) * [GitHub Action](#github-action) * [Contributions](#contributions) * [License](#license) * [Usage](#usage) * [Caching](#caching) * [GitHub Action](#github-action-1) * [Docker image](#docker-image) * [Built-in analyses](#built-in-analyses) * [Project burndown](#project-burndown) * [Files](#files) * [People](#people) * [Churn matrix](#overwrites-matrix) * [Code ownership](#code-ownership) * [Couples](#couples) * [Structural hotness](#structural-hotness) * [Aligned commit series](#aligned-commit-series) * [Added vs changed lines through time](#added-vs-changed-lines-through-time) * [Efforts through time](#efforts-through-time) * [Sentiment (positive and negative comments)](#sentiment-positive-and-negative-comments) * [Everything in a single pass](#everything-in-a-single-pass) * [Plugins](#plugins) * [Merging](#merging) * [Bad unicode errors](#bad-unicode-errors) * [Plotting](#plotting) * [Custom plotting backend](#custom-plotting-backend) * [Caveats](#caveats) * [Burndown Out-Of-Memory](#burndown-out-of-memory) * [Roadmap](#roadmap) ## Overview Hercules is an amazingly fast and highly customizable Git repository analysis engine written in Go. Batteries are included. Powered by [go-git](https://github.com/go-git/go-git). *Notice (November 2020): the main author is back from the limbo and is gradually resuming the development. See the [roadmap](#roadmap).* There are two command-line tools: `hercules` and `labours`. The first is a program written in Go which takes a Git repository and executes a Directed Acyclic Graph (DAG) of [analysis tasks](doc/PIPELINE_ITEMS.md) over the full commit history. The second is a Python script which shows some predefined plots over the collected data. These two tools are normally used together through a pipe. It is possible to write custom analyses using the plugin system. It is also possible to merge several analysis results together - relevant for organizations. The analyzed commit history includes branches, merges, etc. Hercules has been successfully used for several internal projects at [source{d}](https://sourced.tech). There are blog posts: [1](https://blog.sourced.tech/post/hercules-v4), [2](https://blog.sourced.tech/post/hercules) and a [presentation](http://vmarkovtsev.github.io/gowayfest-2018-minsk/). Please [contribute](#contributions) by testing, fixing bugs, adding [new analyses](https://github.com/src-d/hercules/issues?q=is%3Aissue+is%3Aopen+label%3Anew-analysis), or coding swagger!

The DAG of burndown and couples analyses with UAST diff refining. Generated with hercules --burndown --burndown-people --couples --feature=uast --dry-run --dump-dag doc/dag.dot https://github.com/src-d/hercules

torvalds/linux line burndown (granularity 30, sampling 30, resampled by year). Generated with hercules --burndown --first-parent --pb https://github.com/torvalds/linux | labours -f pb -m burndown-project in 1h 40min.

## Installation Grab `hercules` binary from the [Releases page](https://github.com/src-d/hercules/releases). `labours` is installable from [PyPi](https://pypi.org/): ``` pip3 install labours ``` [`pip3`](https://pip.pypa.io/en/stable/installing/) is the Python package manager. Numpy and Scipy can be installed on Windows using http://www.lfd.uci.edu/~gohlke/pythonlibs/ ### Build from source You are going to need Go (>= v1.11) and [`protoc`](https://github.com/google/protobuf/releases). ``` git clone https://github.com/src-d/hercules && cd hercules make pip3 install -e ./python ``` ### GitHub Action It is possible to run Hercules as a [GitHub Action](https://help.github.com/en/articles/about-github-actions): [Hercules on GitHub Marketplace](https://github.com/marketplace/actions/hercules-insights). Please refer to the [sample workflow](.github/workflows/main.yml) which demonstrates how to setup. ## Contributions ...are welcome! See [CONTRIBUTING](CONTRIBUTING.md) and [code of conduct](CODE_OF_CONDUCT.md). ## License [Apache 2.0](LICENSE.md) ## Usage The most useful and reliably up-to-date command line reference: ``` hercules --help ``` Some examples: ``` … ``` `labours -i /path/to/yaml` allows to read the output from `hercules` which was saved on disk. ### Caching It is possible to store the cloned repository on disk. The subsequent analysis can run on the corresponding directory instead of cloning from scratch: ``` # First time - cache hercules https://github.com/git/git /tmp/repo-cache # Second time - use the cache hercules --some-analysis /tmp/repo-cache ``` ### GitHub Action The action produces the artifact named `hercules_charts`. Since it is currently impossible to pack several files in one artifact, all the charts and Tensorflow Projector files are packed in the inner tar archive. In order to view the embeddings, go to [projector.tensorflow.org](https://projector.tensorflow.org), click "Load" and choose the two TSVs. Then use UMAP or T-SNE. ### Docker image ``` docker run --rm srcd/hercules hercules --burndown --pb https://github.com/git/git | docker run --rm -i -v $(pwd):/io srcd/hercules labours -f pb -m burndown-project -o /io/git_git.png ``` ### Built-in analyses #### Project burndown ``` hercules --burndown labours -m burndown-project ``` Line burndown statistics for the whole repository. Exactly the same what [git-of-theseus](https://github.com/erikbern/git-of-theseus) does but much faster. Blaming is performed efficiently and incrementally using a custom RB tree tracking algorithm, and only the last modification date is recorded while running the analysis. All burndown analyses depend on the values of *granularity* and *sampling*. Granularity is the number of days each band in the stack consists of. Sampling is the frequency with which the burnout state is snapshotted. The smaller the value, the more smooth is the plot but the more work is done. There is an option to resample the bands inside `labours`, so that you can define a very precise distribution and visualize it different ways. Besides, resampling aligns the bands across periodic boundaries, e.g. months or years. Unresampled bands are apparently not aligned and start from the project's birth date. #### Files ``` hercules --burndown --burndown-files labours -m burndown-file ``` Burndown statistics for every file in the repository which is alive in the latest revision. Note: it will generate separate graph for every file. You don't want to run it on repository with many files. #### People ``` hercules --burndown --burndown-people [--people-dict=/path/to/identities] labours -m burndown-person ``` Burndown statistics for the repository's contributors. If `--people-dict` is not specified, the identities are discovered by the following algorithm: 0. We start from the root commit towards the HEAD. Emails and names are converted to lower case. 1. If we process an unknown email and name, record them as a new developer. 2. If we process a known email but unknown name, match to the developer with the matching email, and add the unknown name to the list of that developer's names. 3. If we process an unknown email but known name, match to the developer with the matching name, and add the unknown email to the list of that developer's emails. If `--people-dict` is specified, it should point to a text file with the custom identities. The format is: every line is a single developer, it contains all the matching emails and names separated by `|`. The case is ignored. #### Overwrites matrix

Wireshark top 20 devs - overwrites matrix

``` hercules --burndown --burndown-people [--people-dict=/path/to/identities] labours -m overwrites-matrix ``` Beside the burndown information, `--burndown-people` collects the added and deleted line statistics per developer. Thus it can be visualized how many lines written by developer A are removed by developer B. This indicates collaboration between people and defines expertise teams. The format is the matrix with N rows and (N+2) columns, where N is the number of developers. 1. First column is the number of lines the developer wrote. 2. Second column is how many lines were written by the developer and deleted by unidentified developers (if `--people-dict` is not specified, it is always 0). 3. The rest of the columns show how many lines were written by the developer and deleted by identified developers. The sequence of developers is stored in `people_sequence` YAML node. #### Code ownership

Ember.js top 20 devs - code ownership

``` hercules --burndown --burndown-people [--people-dict=/path/to/identities] labours -m ownership ``` `--burndown-people` also allows to draw the code share through time stacked area plot. That is, how many lines are alive at the sampled moments in time for each identified developer. #### Couples

torvalds/linux files' coupling in Tensorflow Projector

``` hercules --couples [--people-dict=/path/to/identities] labours -m couples -o [--couples-tmp-dir=/tmp] ``` **Important**: it requires Tensorflow to be installed, please follow [official instructions](https://www.tensorflow.org/install/). The files are coupled if they are changed in the same commit. The developers are coupled if they change the same file. `hercules` records the number of couples throughout the whole commit history and outputs the two corresponding co-occurrence matrices. `labours` then trains [Swivel embeddings](https://github.com/src-d/tensorflow-swivel) - dense vectors which reflect the co-occurrence probability through the Euclidean distance. The training requires a working [Tensorflow](http://tensorflow.org) installation. The intermediate files are stored in the system temporary directory or `--couples-tmp-dir` if it is specified. The trained embeddings are written to the current working directory with the name depending on `-o`. The output format is TSV and matches [Tensorflow Projector](http://projector.tensorflow.org/) so that the files and people can be visualized with t-SNE implemented in TF Projector. #### Structural hotness ``` 46 jinja2/compiler.py:visit_Template [FunctionDef] 42 jinja2/compiler.py:visit_For [FunctionDef] 34 jinja2/compiler.py:visit_Output [FunctionDef] 29 jinja2/environment.py:compile [FunctionDef] 27 jinja2/compiler.py:visit_Include [FunctionDef] 22 jinja2/compiler.py:visit_Macro [FunctionDef] 22 jinja2/compiler.py:visit_FromImport [FunctionDef] 21 jinja2/compiler.py:visit_Filter [FunctionDef] 21 jinja2/runtime.py:__call__ [FunctionDef] 20 jinja2/compiler.py:visit_Block [FunctionDef] ``` Thanks to Babelfish, hercules is able to measure how many times each structural unit has been modified. By defa

Issues· 0 open

View all issuesOpen on GitHub

No open issues yet, or sync has not completed.

> Tags

Goburndowngitgit-analysismachine-learning

No comments yet. Be the first to share.

> Details

PublishedAug 1, 2026
UpdatedSep 17, 2026
Category开发工具
PricingOpen source

> Related tools

V
VS Code
流行的开源代码编辑器
G
Git
分布式版本控制系统
V
Vite
下一代前端构建工具