- Go 97.8%
- Nix 2.2%
| cmd/gettit | ||
| .gitignore | ||
| devenv.lock | ||
| devenv.nix | ||
| devenv.yaml | ||
| go.mod | ||
| go.sum | ||
| LICENSE | ||
| README.md | ||
gettit
gettit is a Go CLI that captures network assets loaded by a webpage and stores them locally for inspection.
It uses a headless browser (chromedp) to discover requests and re-fetches responses via HTTP for reliability.
What it does
- loads a webpage in a real browser context
- observes network requests (JS, XHR, CSS, images, etc.)
- re-fetches responses via HTTP
- writes files to disk
- records full metadata (including query parameters) in an index
This tool is intended for analysis, not for recreating a working offline site.
Installation
Requirements:
- Go
- Chrome or Chromium
Install dependencies:
go mod tidy
Build the binary:
go build -o gettit ./cmd/gettit
If Chrome is not detected automatically:
export CHROME_PATH=/path/to/chrome
Usage
./gettit https://example.com
Optional output directory:
./gettit https://example.com output_dir
With flags:
./gettit --only js,xhr --same-domain --timeout 15 https://example.com
With flags and custom output directory:
./gettit --only js,xhr https://example.com output_dir
Flags
--only js,xhr Only store selected resource types
--same-domain Ignore third-party domains
--timeout 12 Maximum capture duration in seconds
Output
By default, data is written per domain to:
dump/<domain>/
Example:
dump/ad.nl/
js/
xhr/
css/
img/
html/
other/
index.json
If an output directory is provided, it is used directly instead:
gettit https://example.com output_dir
→ writes to:
output_dir/
If the output directory already exists, it is removed before a new run.
Categorization is best-effort based on MIME type.
Filenames
Format:
<domain>__<shortpath>__<hash>.<ext>
Example:
ad.nl__ssosession__63a011.json
Details:
- subdomains are removed (
api.foo.example.com→example.com) - only the last 1–2 path segments are used
- query parameters are not included
- filenames are kept short for navigation
- uniqueness is ensured by hashing the full URL (including query parameters)
index.json
Contains metadata for each file:
{
"<filename>": {
"url": "...",
"method": "...",
"status": 200,
"mime": "...",
"content_type": "...",
"size": 1234,
"resource_type": "...",
"headers": { ... },
"saved_path": "...",
"hash": "...",
"time": "..."
}
}
Notes:
urlcontains the full request (including query parameters)hashis derived from the response body- a subset of request headers is stored
Behavior
- identical responses are deduplicated by content hash
- performs a simple scroll to trigger lazy-loaded content
- stops early when the network is idle (≈1.5s without activity)
- always stops after
--timeoutif activity continues - shows a live progress bar during capture
- duplicate URLs are only fetched once
Limitations
- idle detection is heuristic; some sites with continuous polling may run until timeout
- no complex interaction (clicks, forms, etc.)
- some requests may be missed if triggered late
- WebSockets and streaming responses are not captured
- authenticated resources may not be accessible
- HTTP re-fetch may differ slightly from browser response
Notes
- filenames are optimized for readability; full details are in
index.json - response bodies are stored as-is