aget is a local-first CLI for turning web pages and explicitly approved
browser/session state into agent-ready content.
The current product surface is the Rust binary and its structured envelope. Normal use does not require external scraper or browser-control tools.
aget get <url>for HTTP(S),raw:,raw://, andfile://inputs.aget current-tab --cdp-port <port> --allow-private-contentfor explicitly approved local Chrome DevTools tab extraction.aget batch <url> [<url>...]for bounded concurrent fetches with per-URL artifacts and a batch manifest.aget map <url>andaget map --artifact <run-id>for one-page link discovery without recursive fetching.aget crawl <url> --limit <n>for same-origin/path-bounded traversal with a manifest and per-page artifacts.aget search-page --artifact <run-id> --query <text>for deterministic local snippet search over existing page artifacts.aget extract --artifact <run-id> ...andaget extract --manifest <path> ...for deterministic structured extraction from existing artifacts.aget interact <url|current-tab> --actions <path>for explicitly approved, file-driven local browser actions with audit, capture, and extract artifacts.aget doctorfor local readiness diagnostics.aget setup-skillsfor installing global skills and project-local agent adapters from the installed binary.aget artifacts ...for listing, inspecting, deleting, and pruning local run artifacts underAGET_HOME.- Opt-in debug artifacts with
--capture-traceand browser-backed--capture-screenshot. aget session ...commands for listing, inspecting, deleting, composing, importing, authorizing, and bootstrapping local sessions.--envelope jsonfor stable agent/tool output.--content-format markdown|html|text|json.--output <path>for writing extracted content to a chosen file.- CSS
--selector,--exclude-selector, and--wait-for-selectorshaping. --max-charsdeterministic post-extraction truncation.--fresh,--cache-policy, and--cache-ttlfor reusable public HTTP(S) fetches.- Repeated
--session <name>for explicit named-session replay and composition. - Replay-time checks that reject sessions outside the requested URL's saved scope.
- Local run artifacts under
~/.aget/runs. - Project skill guidance in
skills/aget/SKILL.md. - Project-local OpenCode tools in
.opencode/tools/aget.ts.
Planned CLI work is tracked in workpads/post-migration/tasks.md.
Developer build from a checkout:
cargo build
cargo run -- --helpDeveloper install from this checkout:
cargo install --path .
aget --helpRegistry install after the crates.io release:
cargo install aget --locked
aget --helpSource install from GitHub:
cargo install --git https://github.com/nicolaslara/aget --tag v0.1.0 --locked aget
aget --helpInstall agent integrations from the installed binary:
aget setup-skills --all --forceThat installs global aget skills for Codex, Claude Code, Gemini CLI, and
Windsurf. To also install project-local adapters for OpenCode, Cursor, and
GitHub Copilot in another repository, pass the target project directory:
aget setup-skills --all --project-dir /path/to/repo --forceFrom a checkout or unpacked release tarball, the helper script provides the same setup flow:
scripts/install-agent-integrations.sh --all --forceWith project-local adapters:
scripts/install-agent-integrations.sh --all --project-dir /path/to/repo --forceThe helper defaults to copying files. Use --symlink from a development
checkout when local integration edits should be reflected after restarting or
reloading the agent harness. The installed aget setup-skills command copies
embedded templates and rejects --symlink; use the checkout script for symlink
mode.
Current release artifact:
dist/aget-v0.1.0-aarch64-apple-darwin.tar.gz
dist/aget-v0.1.0-aarch64-apple-darwin.tar.gz.sha256
dist/SHA256SUMS
Install the CLI from a release tarball:
curl -fLO https://github.com/nicolaslara/aget/releases/download/v0.1.0/aget-v0.1.0-aarch64-apple-darwin.tar.gz
curl -fLO https://github.com/nicolaslara/aget/releases/download/v0.1.0/aget-v0.1.0-aarch64-apple-darwin.tar.gz.sha256
shasum -a 256 -c aget-v0.1.0-aarch64-apple-darwin.tar.gz.sha256
tar -xzf aget-v0.1.0-aarch64-apple-darwin.tar.gz
install -d "$HOME/.local/bin"
install -m 755 aget-v0.1.0-aarch64-apple-darwin/aget "$HOME/.local/bin/aget"
"$HOME/.local/bin/aget" --helpInstall only the Codex skill globally from a checkout:
scripts/install-codex-skill.sh --symlink --forceInstall only the Codex skill globally from the release tarball:
CODEX_HOME="${CODEX_HOME:-$HOME/.codex}"
mkdir -p "$CODEX_HOME/skills"
rm -rf "$CODEX_HOME/skills/aget"
cp -R aget-v0.1.0-aarch64-apple-darwin/skills/aget "$CODEX_HOME/skills/aget"The helper defaults to copying the skill. Use --symlink from a development
checkout when you want local skill edits to be picked up after a Codex restart.
Restart Codex after installing or replacing a global skill. Release artifact
naming and smoke gates are tracked in workpads/post-migration/release-plan.md.
Prerequisites:
- Rust and Cargo.
- Local Chrome or Chromium when using JavaScript-rendered pages, Chrome import, login flows, or current-tab extraction.
- Optional
cmuxonly foraget session import cmux.
Fetch a public page as markdown:
aget get https://example.comFetch with a structured envelope:
aget --envelope json get https://example.comWrite content to a file:
aget get https://example.com --output /tmp/example.mdShape output:
aget --envelope json get https://example.com \
--content-format markdown \
--selector main \
--max-chars 4000Extract an explicitly approved current tab:
aget --envelope json current-tab \
--cdp-port 9222 \
--allow-private-content \
--output /tmp/current-tab.mdFetch several pages into a manifest:
aget --envelope json batch https://example.com https://example.org \
--concurrency 2 \
--output-dir /tmp/aget-batchDiscover links from one page without crawling them:
aget --envelope json map https://example.com/docs \
--same-origin \
--same-path \
--max-links 200Crawl a bounded docs section:
aget --envelope json crawl https://example.com/docs/ \
--limit 25 \
--max-depth 2 \
--concurrency 2Search a prior page artifact:
aget --envelope json search-page --artifact run-123 --query "install instructions"Extract structured values from prior artifacts:
aget --envelope json extract --artifact run-123 --field tables --field links
aget --envelope json extract --manifest /tmp/aget-crawl/manifest.json --field headingsValidate a browser action plan without executing it:
aget --envelope json interact https://example.com/app \
--actions /tmp/actions.json \
--allow-actions \
--dry-runRun approved actions against an explicitly selected current tab:
aget --envelope json interact current-tab \
--cdp-port 9222 \
--actions /tmp/actions.json \
--allow-actions \
--allow-private-contentTop-level commands:
aget get <url>
aget batch <url> [<url>...]
aget map <url>
aget map --artifact <run-id>
aget crawl <url> --limit <n>
aget search-page --artifact <run-id> --query <text>
aget extract --artifact <run-id> --field <name>
aget extract --manifest <path> --schema <path>
aget interact <url|current-tab> --actions <path> --allow-actions
aget current-tab --cdp-port <port> --allow-private-content
aget session <command>
aget artifacts <command>
aget doctor
aget setup-skills --all [--project-dir <dir>]
aget get:
aget get <url>
[--session <name>...]
[--envelope <json|none>]
[--output <path>]
[--timeout <seconds>]
[--content-format <markdown|html|text|json>]
[--inline-content <auto|always|never>]
[--selector <css>]
[--exclude-selector <css>]
[--wait-for-selector <css>]
[--max-chars <n>]
[--fresh]
[--cache-policy <auto|refresh|off>]
[--cache-ttl <seconds>]
[--capture-trace]
[--capture-screenshot]
[--backend-option <aget.key=value>...]
--capture-trace writes a redacted debug-trace.json under the run directory.
--capture-screenshot writes screenshot.png only when the request uses a
browser/CDP rendering path; static extraction emits a warning instead of
capturing visual content.
Cache reuse applies only to eligible public HTTP(S) fetches. Session-backed,
current-tab, raw:, and file:// inputs record cache metadata but do not
read or write reusable cache entries.
aget batch:
aget batch <url> [<url>...]
[--file <path>]
[--stdin]
[--session <name>...]
[--envelope <json|none>]
[--output-dir <path>]
[--timeout <seconds>]
[--content-format <markdown|html|text|json>]
[--selector <css>]
[--exclude-selector <css>]
[--wait-for-selector <css>]
[--max-chars <n>]
[--fresh]
[--cache-policy <auto|refresh|off>]
[--cache-ttl <seconds>]
[--backend-option <aget.key=value>...]
[--concurrency <1-8>]
[--fail-fast]
aget batch accepts positional URLs, newline-delimited --file input, or
newline-delimited --stdin input. It writes manifest.json, manifest.md,
and per-item content/metadata artifacts under --output-dir; without
--output-dir, it creates a local batch run directory under AGET_HOME.
Duplicate inputs are reported as skipped. Partial failures are reported in the
manifest and return a non-zero exit code.
aget map:
aget map <url>
aget map --artifact <run-id>
[--session <name>...]
[--envelope <json|none>]
[--selector <css>]
[--exclude-selector <css>]
[--wait-for-selector <css>]
[--same-origin|--any-origin]
[--same-path|--any-path]
[--include <pattern>...]
[--exclude <pattern>...]
[--max-links <n>]
[--fresh]
[--cache-policy <auto|refresh|off>]
[--cache-ttl <seconds>]
[--content-type <type>...]
[--output <markdown|json>]
[--backend-option <aget.key=value>...]
aget map fetches one page as HTML or reads one successful internal get run
artifact, extracts links, deduplicates normalized absolute URLs, drops
fragments, and emits links.json plus links.md under a local AGET_HOME run
directory. It does not fetch discovered links. By default it keeps only links
on the same origin and same path prefix as the source; use --any-origin or
--any-path to widen discovery. Artifact mode rejects caller-owned external
--output content paths.
aget crawl:
aget crawl <url> --limit <n>
[--max-depth <n>]
[--concurrency <1-8>]
[--same-origin|--any-origin]
[--same-path|--any-path]
[--allow-domain <domain>...]
[--include <pattern>...]
[--exclude <pattern>...]
[--session <name>...]
[--envelope <json|none>]
[--output-dir <path>]
[--content-format <markdown|html|text|json>]
[--selector <css>]
[--exclude-selector <css>]
[--wait-for-selector <css>]
[--max-chars <n>]
[--fresh]
[--cache-policy <auto|refresh|off>]
[--cache-ttl <seconds>]
[--backend-option <aget.key=value>...]
aget crawl requires --limit and caps v1 crawls at 100 pages, depth 5, and
concurrency 8. It fetches pages through the same get pipeline, discovers
links from fetched HTML, and writes manifest.json, manifest.md, and
per-page artifacts under --output-dir or a local AGET_HOME run directory.
Traversal defaults to the start URL's origin and path prefix; --any-origin,
--any-path, and --allow-domain must be explicit. Crawl records partial
failures in the manifest and returns a non-zero exit code when any page fails.
aget search-page:
aget search-page --artifact <run-id> --query <text>
[--max-results <n>]
[--context-chars <n>]
[--allow-private-content]
[--output <markdown|json>]
[--envelope <json|none>]
aget search-page reads an existing successful get artifact, scores local
sections by heading and keyword matches, and writes search-results.json plus
search-results.md under a fresh local run directory. Sensitive artifacts
require --allow-private-content before snippets are emitted.
aget extract:
aget extract --artifact <run-id>
aget extract --manifest <path>
[--schema <path>]
[--field <headings|links|tables|definitions|metadata>...]
[--allow-private-content]
[--output <json|markdown>]
[--envelope <json|none>]
aget extract reads existing successful get artifacts or successful items in
a batch/crawl manifest. Built-in fields extract headings, links, tables,
definition lists, and metadata without refetching or calling an LLM. Schema
files can add selector fields for HTML artifacts and json fields for JSON
content paths. Sensitive source artifacts require --allow-private-content
before structured values are emitted. Results are written to
extract-results.json, extract-results.md, and metadata.json under a fresh
local run directory. Artifact mode reads only internal run content; it refuses
caller-owned external --output paths. Manifest mode requires each item to
carry matching aget item metadata and keeps content reads inside the manifest
directory.
aget interact:
aget interact <url>
aget interact current-tab --cdp-port <port>
--actions <path>
--allow-actions
[--allow-private-content]
[--allow-sensitive-input]
[--allow-submit]
[--capture-screenshot]
[--session <name>...]
[--dry-run]
[--timeout <seconds>]
[--envelope <json|none>]
aget interact executes a local JSON action plan and always writes redacted
actions-request.json, actions-result.json, and metadata.json under a run
directory. Non-dry-run execution currently supports URL and current-tab sources
without named session replay; session-backed action execution returns a
structured action_not_supported result until session replay is wired into the
interact executor.
Action plans use schema version aget.actions.v1:
{
"schema_version": "aget.actions.v1",
"actions": [
{"type": "wait", "selector": "main"},
{"type": "click", "selector": "button.next"},
{"type": "capture", "name": "after-click", "html": true},
{
"type": "extract",
"name": "main",
"selector": "main",
"content_format": "markdown"
}
]
}Supported action types are wait, click, type, select, submit,
capture, and extract. --allow-actions is always required because plans can
mutate browser state. current-tab and session-backed plans require
--allow-private-content; sensitive typed input requires
--allow-sensitive-input; submit requires both confirm: true in the action
and --allow-submit; screenshot captures require --capture-screenshot.
Selectors are strict by default. Mutation actions require one matching element.
Read actions such as wait, capture, and extract may opt into
"match": "first"; otherwise zero or ambiguous matches fail. capture writes
HTML and screenshot files under artifacts.captures; selector-backed
screenshots capture the selected element region. extract runs the owned
extraction pipeline against the current DOM and writes output under
artifacts.extracts. Sensitive interact runs redact URL query strings and
fragments in envelopes and metadata.
aget current-tab:
aget current-tab --cdp-port <port>
--allow-private-content
[--envelope <json|none>]
[--output <path>]
[--timeout <seconds>]
[--content-format <markdown|html|text|json>]
[--inline-content <auto|always|never>]
[--selector <css>]
[--exclude-selector <css>]
[--wait-for-selector <css>]
[--max-chars <n>]
[--capture-trace]
[--capture-screenshot]
[--backend-option <aget.key=value>...]
Session commands:
aget session list
aget session inspect <name>
aget session delete <name>
aget session compose <new-name> --session <name> [--session <name>...]
aget session authorize <name> --url <url> --browser chrome --browser-profile <profile> --allow-domain <domain>
aget session import cmux --surface <surface> --name <name> --allow-domain <domain>
aget session import browser --browser chrome --browser-profile <profile> --name <name> --allow-domain <domain>
aget session import chrome --chrome-profile <profile> --name <name> --allow-domain <domain>
aget session login start <name> --url <login-or-target-url> [--session <provider>...]
aget session login finish <name>
aget session login cancel <name>
Artifact commands:
aget artifacts list
aget artifacts inspect <run-id>
aget artifacts delete <run-id> --yes
aget artifacts prune --older-than <duration> [--keep-last <n>] [--dry-run|--yes]
aget artifacts prune --keep-last <n> [--dry-run|--yes]
aget artifacts prune --max-bytes <bytes> [--keep-last <n>] [--dry-run|--yes]
Run aget <command> --help for full option details.
By default, aget get prints extracted page content directly. Markdown is the
default content format.
--envelope json prints a stable control-plane response:
{
"ok": true,
"schema_version": "aget.envelope.v1",
"command": "get",
"data": {
"url": "https://example.com",
"final_url": "https://example.com/",
"content_format": "markdown",
"content": "# Example\n...",
"artifacts": {
"content": "/Users/me/.aget/runs/abc123/output.md",
"metadata": "/Users/me/.aget/runs/abc123/metadata.json",
"debug": {
"trace": {
"path": "/Users/me/.aget/runs/abc123/debug-trace.json",
"media_type": "application/json",
"sensitive": false
}
}
},
"cache": {
"status": "miss",
"policy": "auto",
"eligible": true,
"key": "0123456789abcdef",
"ttl_seconds": 86400,
"age_seconds": null,
"reason": "cache entry not found"
},
"usage": {
"fetched_bytes": 12000,
"content_bytes": 2400,
"estimated_tokens": 600,
"estimated_tokens_saved": 2400
},
"sessions": [],
"sensitive": false
},
"warnings": [],
"timing_ms": {
"total": 482
}
}Error envelopes use the same schema:
{
"ok": false,
"schema_version": "aget.envelope.v1",
"command": "get",
"error": {
"code": "extraction_failed",
"message": "aget extraction failed"
}
}--inline-content auto includes data.content for non-sensitive fetches and
omits it for session-backed or current-tab output. Use --output and read the
artifact path for large or private content. Use --inline-content always only
when private content should be embedded in the JSON envelope.
data.cache.status is one of miss, hit, stale, refresh, disabled,
or ineligible. Cache entries live under AGET_HOME/cache, are keyed by URL
and extraction-shaping options, and store the full untruncated extracted
content so each run can still apply its own --max-chars. data.usage
reports fetched bytes when known, final content bytes, and rough token
estimates for agent-side budgeting.
Debug artifacts are never captured by default. debug-trace.json records
control-plane diagnostics such as options, cache state, warnings, timings, and
errors; it does not include extracted page content. Screenshot artifacts can
contain visual page content and are marked sensitive when the source is
session-backed or current-tab.
aget starts with an empty session by default. Authenticated state is used only
when the caller passes explicit local session names.
Import a scoped browser session from an approved Chrome profile:
aget --envelope json session import browser \
--browser chrome \
--browser-profile Default \
--name workdocs \
--allow-domain docs.example.comFetch with the imported session:
aget --envelope json get "https://docs.example.com/account" \
--session workdocs \
--content-format markdown \
--output /tmp/workdocs.mdImport and verify in one flow:
aget --envelope json session authorize workdocs \
--url "https://docs.example.com/account" \
--browser chrome \
--browser-profile Default \
--allow-domain docs.example.com \
--must-contain "Account" \
--output /tmp/workdocs-check.mdStart a controlled user-driven login flow:
aget --envelope json session login start workdocs \
--url "https://docs.example.com/login"
# user completes login in the opened browser
aget --envelope json session login finish workdocsFor OAuth-backed sites, a reusable provider session may be injected into the controlled login browser:
aget --envelope json session login start workdocs \
--url "https://docs.example.com/login" \
--session oauth
aget --envelope json session login finish workdocsProvider sessions are for login bootstrap only. Do not pass a provider session
to aget get for target-site content. Replay-scope checks are a guardrail that
reject out-of-scope session use; they do not turn provider credentials into
target-site authorization. Fetch target content with the relying-party session
saved by login finish.
Chrome/CDP is the verified browser path today.
aget session import browser --browser chrome is the browser-neutral command
for the verified Chrome path. Other browser names currently return structured
unsupported or usage errors until they are proven safe.
Prefer named Chrome profiles through --browser-profile. Explicit
--profile-path is advanced: it may launch that local profile directory
directly, so use it only with a disposable or explicitly approved profile path.
Import stores only the scoped exported cookies/storage in the named aget
session. It does not copy, retain, or manage the original browser profile as
part of aget state. For login flows, the default aget-owned temporary profile
is cleaned up after finish or cancel; a caller-provided custom profile path is
caller-owned and is not deleted or pruned by aget.
current-tab never scans profiles or common ports. The caller must provide an
explicit local CDP port and --allow-private-content.
- Only process content the user is authorized to access.
agetdoes not collect, script, or store credentials.- Session replay is opt-in per request with
--session <name>. - Session replay is rejected when selected sessions are outside the request host scope.
- Run artifacts live under
~/.aget/runs. - Debug traces and screenshots are opt-in run artifacts. Treat screenshots as page content; for sensitive sources, inspect only the local path needed for the task and avoid embedding visual/private content into model context.
artifacts deleteandartifacts pruneremove only internal run directories underAGET_HOME/runs; caller-provided--outputfiles outside those run directories are preserved.session deleteremoves saved auth/session state only. It does not delete previous extracted content artifacts or caller-provided--outputfiles.- Imported browser cookies and localStorage are credential-equivalent bearer
material.
session inspectredacts values by default. current-tabmay read private authenticated browser content from the selected tab. Prefer--outputand avoid--inline-content alwaysunless explicitly requested.- Do not use
agetto bypass access controls, paywalls, or site policies.
aget is a generic fetcher. It does not contain site-specific paywall, login,
rate-limit, or access-state detection. The calling agent interprets returned
content and decides whether to ask the user for a session, a narrower selector,
or a different URL.
This repo includes project-local OpenCode tools in .opencode/tools/aget.ts.
They call the installed aget binary with --envelope json and return the same
structured envelopes as terminal usage.
Available tools:
aget_fetchaget_batchaget_mapaget_crawlaget_search_pageaget_extractaget_interactaget_artifacts_listaget_artifacts_inspectaget_doctoraget_session_listaget_session_inspectaget_session_import_chrome
These wrappers are thin CLI/envelope callers. They do not start a server, daemon, MCP layer, or direct Rust API.
Use AGET_OPENCODE_BIN to point OpenCode at a development binary:
cargo build
AGET_OPENCODE_BIN="$PWD/target/debug/aget" opencodecargo fmt --check
cargo test
git diff --check
bash -n scripts/demo_real_cli.shTesting layers:
- Focused unit and integration tests for parser, envelope, extraction, session, subprocess, artifact, and redaction behavior.
tests/mock_site_cli.rsandtests/support/mock_site.rsfor deterministic local auth/session coverage without real credentials or live sites.- Behavior-level parity tests for historical source projects only where
agetimplements the corresponding feature. These tests should be local, license-aware, and adapted toagetbehavior rather than copied blindly. - Ignored/manual real-browser checks for confidence in user-authorized local Chrome flows.
Early aget prototypes used external source projects as implementation and
behavior references. The current default runtime is owned Rust code. Historical
source snapshots remain useful for parity research and license-aware behavior
comparison, but they are not normal user prerequisites.
The package declares MIT in Cargo.toml. This repository does not currently
include a separate LICENSE file.