ALIENS EYE
AI-OSINT Username Scanner
Advanced AI-Powered Social Media Username Finder
Scan 840+ platforms with ML-blended detection
## Highlights
- **840+ platforms** scanned asynchronously in seconds
- **ML + heuristic detection** — a trained model blended with 30 structural signals (HTTP status, DOM shape, keywords, fingerprints) instead of naive status-code checks
- **Profile extraction** — display name, bio, and avatar pulled from each hit (OpenGraph / JSON-LD / per-site CSS)
- **Cross-site correlation** — cluster profiles that look like the same person by avatar hash, bio, shared links, and name (`--correlate`)
- **Recursive expansion** — follow linked usernames out of bios and re-scan them (`--recurse-depth N`)
- **Domain check** — is `.{com,io,net,…}` registered and live? (`--domains`)
- **Watch mode** — re-scan on an interval and alert on changes, optionally to a webhook (`--watch 6h --notify `)
- **Resumable scans** — checkpoint progress and continue after an interruption (`--resume file.jsonl`)
- **Modern terminal UI** — live progress, sorted result tables, summary panels (powered by [rich](https://github.com/Textualize/rich)); plus an interactive browser (`aliens_eye tui`, optional extra)
- **MCP server** — expose scanning to LLM agents (`aliens_eye serve`, optional extra)
- **Proxy & Tor support** — `--proxy socks5://...` or just `--tor`
- **Site filtering** — `--site github,reddit`, `--exclude-site`, `--no-nsfw`, plus drop-in `sites.d/` plugin site maps
- **Calibrated self-check** — `aliens_eye selfcheck` reports precision / recall / F1 / FPR per site
- **Reproducible evaluation** — record a frozen response corpus once, then replay it for identical metrics run to run (`aliens_eye corpus record` / `selfcheck --corpus`)
- **Ablations and baselines** — `aliens_eye eval ablate` scores detector configurations with bootstrap confidence intervals; `eval external` compares against Sherlock / Maigret / WhatsMyName rules on the same stored responses
- **Retrainable + active learning** — retrain with `aliens_eye train`, or hand-label uncertain hits with `aliens_eye label`
- **Reports** in JSON, CSV, HTML, Markdown, PDF, and graph formats (GEXF, Mermaid, Maltego CSV)
- **Playwright fallback** for JavaScript-heavy pages (optional extra)
## Install
```bash
pip install aliens-eye
```
Optional extras:
```bash
pip install "aliens-eye[browser]" # Playwright fallback for hard pages
python -m playwright install chromium
pip install "aliens-eye[train]" # scikit-learn, for retraining the ML model
pip install "aliens-eye[correlate]" # Pillow, for avatar-image matching in --correlate
pip install "aliens-eye[pdf]" # reportlab, for --format pdf
pip install "aliens-eye[tui]" # textual, for the interactive `tui` browser
pip install "aliens-eye[serve]" # mcp, for the `serve` MCP server
```
Or with Docker:
```bash
docker build -t aliens-eye .
docker run --rm -it aliens-eye username
```
From source:
```bash
git clone https://github.com/arxhr007/Aliens_eye.git
cd Aliens_eye
pip install -e .
```
## Usage
```
…
```
Custom platforms: drop a `{ "site_name": "https://site/{}" }` JSON file into
`./sites.d/` (or the user config dir's `sites.d/`) and it is merged automatically;
`--sites-dir DIR` adds another location.
## How detection works
Every response is converted into a 30-dimensional feature vector: HTTP status buckets, username placement (path/title/meta/canonical), error and profile keywords, DOM structure (images, forms, profile/error CSS classes), structured-data signals (og:type, JSON-LD Person), response timing, redirect counts, and per-site fingerprint matches learned from previous scans.
Two judges then vote:
1. **Heuristic engine** — weighted scoring over the features
2. **ML model** — logistic regression trained on labeled scans of real (and deliberately fake) accounts, shipped with the package and running in pure Python (no sklearn needed at runtime)
The blended probability maps to **Found / Maybe / Not Found** with a confidence percentage. The loaded model supplies both the blend weight and the thresholds — the shipped model uses `0.6 * ml + 0.4 * heuristic`, Found above `0.556`, Not Found below `0.322`. If a model file is missing or invalid, the scanner silently falls back to heuristics with the defaults in `core/detector.py` (`0.4` ML weight, `0.6` / `0.35` thresholds). See [WORKING.md](WORKING.md) for the full table.
> **Detection accuracy is preliminary.** The shipped model was fit on 368 samples from 43 platforms (`cv_f1 = 0.5622`), with ground-truth accounts skewed toward high-profile users. Treat Found/Maybe as leads to verify, not as findings.
### Retraining the model
```bash
pip install "aliens-eye[train]"
# 1. Scan ground-truth accounts + random non-existent usernames to build a dataset
# (reads the train split only; the eval holdout is never touched)
aliens_eye train collect --out dataset.csv --negatives 4
# 2. Fit and export the model
aliens_eye train fit --data dataset.csv --out model.json
# 3. Score it on platforms it never trained on
aliens_eye selfcheck --split holdout --model model.json --report json
# 4. Use it
aliens_eye username --model model.json
```
Ground truth is split **site-disjoint** into `data/selfcheck.json` (train, 30
sites) and `data/eval_holdout.json` (holdout, 13 sites). Scoring `--split train`
measures fit, not generalization, and will read high.
## Configuration
Aliens Eye merges a JSON config file with CLI flags (CLI wins). Search order without `--config`: `./config.json`, then the platform config dir (e.g. `~/.config/aliens_eye/config.json` on Linux, `%LOCALAPPDATA%\aliens_eye` on Windows).
```json
{
"concurrent": 50,
"timeout": 10.0,
"retries": 2,
"rate_limit_delay": 0.2,
"output_dir": "results",
"output_formats": ["json", "csv", "html", "md"],
"use_playwright": false,
"proxy": null,
"use_ml": true,
"exclude_nsfw": false,
"level": "basic"
}
```
## Outputs
Results are saved with timestamped filenames:
- `username_level_YYYYMMDD_HHMMSS.json` — full detail including per-site feature analysis
- `.csv` — flat rows for spreadsheets
- `.html` — styled standalone report
- `.md` — Markdown summary of Found/Maybe hits
## Architecture
The package lives under `src/aliens_eye/`: `core/` (scanner, detector, analyzer, http, exporter, fingerprints), `ml/` (inference, training, dataset collection), `utils/` (rich console layer), and `data/` (sites.json, trained model, ground-truth sets). For internals and flowcharts, see [WORKING.md](WORKING.md).
## Contributing
Issues and PRs welcome — adding sites to `src/aliens_eye/data/sites.json`, expanding the ground-truth set in `selfcheck.json`, or improving the model all directly improve detection. Run `pytest` and `ruff check src tests` before submitting.
## Disclaimer
This tool is for educational purposes and legitimate OSINT research only. You are responsible for complying with laws and site terms of service.