Forensic copy of loaded webfiles using real Chrome browser
  • Go 97.8%
  • Nix 2.2%
Find a file
2026-05-14 15:57:09 +02:00
cmd/gettit fix: detect HTML extension using response content and headers; avoid .bin for HTML pages 2026-05-13 17:10:05 +02:00
.gitignore update gitignore 2026-05-14 15:40:46 +02:00
devenv.lock Init 2026-05-13 15:33:06 +02:00
devenv.nix Update readme for building 2026-05-13 17:04:43 +02:00
devenv.yaml Init 2026-05-13 15:33:06 +02:00
go.mod refactor: split into internal files (capture, fetch, index, naming, filter, progress) and cmd entry; integrate progressbar 2026-05-13 16:16:14 +02:00
go.sum refactor: split into internal files (capture, fetch, index, naming, filter, progress) and cmd entry; integrate progressbar 2026-05-13 16:16:14 +02:00
LICENSE Add license 2026-05-14 15:57:09 +02:00
README.md Update readme for building 2026-05-13 17:04:43 +02:00

gettit

gettit is a Go CLI that captures network assets loaded by a webpage and stores them locally for inspection.

It uses a headless browser (chromedp) to discover requests and re-fetches responses via HTTP for reliability.


What it does

  • loads a webpage in a real browser context
  • observes network requests (JS, XHR, CSS, images, etc.)
  • re-fetches responses via HTTP
  • writes files to disk
  • records full metadata (including query parameters) in an index

This tool is intended for analysis, not for recreating a working offline site.


Installation

Requirements:

  • Go
  • Chrome or Chromium

Install dependencies:

go mod tidy

Build the binary:

go build -o gettit ./cmd/gettit

If Chrome is not detected automatically:

export CHROME_PATH=/path/to/chrome

Usage

./gettit https://example.com

Optional output directory:

./gettit https://example.com output_dir

With flags:

./gettit --only js,xhr --same-domain --timeout 15 https://example.com

With flags and custom output directory:

./gettit --only js,xhr https://example.com output_dir

Flags

--only js,xhr        Only store selected resource types
--same-domain        Ignore third-party domains
--timeout 12         Maximum capture duration in seconds

Output

By default, data is written per domain to:

dump/<domain>/

Example:

dump/ad.nl/
  js/
  xhr/
  css/
  img/
  html/
  other/
  index.json

If an output directory is provided, it is used directly instead:

gettit https://example.com output_dir

→ writes to:

output_dir/

If the output directory already exists, it is removed before a new run.

Categorization is best-effort based on MIME type.


Filenames

Format:

<domain>__<shortpath>__<hash>.<ext>

Example:

ad.nl__ssosession__63a011.json

Details:

  • subdomains are removed (api.foo.example.com → example.com)
  • only the last 1–2 path segments are used
  • query parameters are not included
  • filenames are kept short for navigation
  • uniqueness is ensured by hashing the full URL (including query parameters)

index.json

Contains metadata for each file:

{
  "<filename>": {
    "url": "...",
    "method": "...",
    "status": 200,
    "mime": "...",
    "content_type": "...",
    "size": 1234,
    "resource_type": "...",
    "headers": { ... },
    "saved_path": "...",
    "hash": "...",
    "time": "..."
  }
}

Notes:

  • url contains the full request (including query parameters)
  • hash is derived from the response body
  • a subset of request headers is stored

Behavior

  • identical responses are deduplicated by content hash
  • performs a simple scroll to trigger lazy-loaded content
  • stops early when the network is idle (≈1.5s without activity)
  • always stops after --timeout if activity continues
  • shows a live progress bar during capture
  • duplicate URLs are only fetched once

Limitations

  • idle detection is heuristic; some sites with continuous polling may run until timeout
  • no complex interaction (clicks, forms, etc.)
  • some requests may be missed if triggered late
  • WebSockets and streaming responses are not captured
  • authenticated resources may not be accessible
  • HTTP re-fetch may differ slightly from browser response

Notes

  • filenames are optimized for readability; full details are in index.json
  • response bodies are stored as-is