Preliminary How-to Setup and Use Guide
An early, deliberately unpolished guide, written on 2026-09-30 from how the code works today. It will grow. Every part carries a status tag, so you can see what is real:
- Available works today and is public.
- Untested exists in our private codebase but has not been run end to end yet.
- Draft partly wired, with known gaps.
- Planned not built yet; described so you know where we are heading.
1. What exists today
| Piece | Status | Where |
|---|---|---|
| Core scanner: context, scan, then skeptical multi-round triage | Available | Open source: github.com/Huge/nano-analyzer |
| GitLab merge-request job that posts one report comment | Untested | Private codebase, shared during early access |
| GitHub Action | Draft | Private codebase, shared during early access |
@budgetscan comment command with budget and deadline caps | Planned | — |
| Prepaid credits and hosted scans on zero-data-retention inference | Planned | — |
The public repository is the open-source core. We also maintain a much more advanced private codebase; the CI integrations below live there.
2. Quick start with the command line Available
You need Python 3.8+ and an API key: OpenAI (the default model is gpt-5.4-nano) or OpenRouter (model names with a slash, such as qwen/qwen3-32b). Optionally install ripgrep (rg) so the triage step can search your code. There are no other dependencies.
git clone https://github.com/Huge/nano-analyzer.git
cd nano-analyzer
# OpenAI models (no slash in the model name)
export OPENAI_API_KEY=sk-...
# or OpenRouter models (slash in the model name)
export OPENROUTER_API_KEY=sk-or-...
# scan a file or a whole directory
python3 scan.py ./path/to/src
Every stage calls an LLM, so cost grows with the amount of code and the number of triage rounds. Start with one directory and see what it costs before pointing it at a whole repository.
Common options
| Option | What it does |
|---|---|
--model | Model for all stages, e.g. --model qwen/qwen3-32b |
--parallel N | Total concurrent API calls |
--triage-rounds N | Skeptical review rounds per finding (default 5) |
--triage-threshold | Lowest severity that gets triaged (default medium) |
--min-confidence 0.7 | Only show findings with at least this confidence |
--repo-dir ./ | Repository root the triage step searches; useful when you scan only a subfolder |
--extensions, --ignore-dirs | Narrow the set of files scanned |
--output-dir | Where results are written (default ~/nano-analyzer-results/<timestamp>/) |
3. Reading the results
<timestamp>/
βββ summary.md # start here
βββ findings/ # findings that survived triage
β βββ VULN-001_<file>.md
βββ triage_survivors.md # summary of validated findings
βββ triage.json # every verdict
βββ triages/ # per-finding reasoning
βββ <file>.md / .json / .context.md # raw output per file
A confidence such as 80% [VVIVVβV] is the share of triage rounds that judged the finding valid; the last letter is the arbiter's final call. Always verify a finding by hand before acting on it.
4. GitLab merge-request integration Untested
A CI job that scans the files changed in a merge request and keeps one report discussion up to date on that MR. It exists in our private codebase and has not been run end to end yet, so expect rough edges. Ask us for the files during early access (see Questions).
- Build and publish the image. The Dockerfile installs git, ripgrep and the optional code-search tools:
docker build -f gitlab/Dockerfile -t registry.gitlab.com/<group>/<project>/nano-analyzer:latest . docker push registry.gitlab.com/<group>/<project>/nano-analyzer:latest - Add the job. Copy
scan.pyand thegitlab/folder into the project you want to scan (the wrapper currently runsscan.pyfrom that project's root; this packaging will probably change), then include the template and set the image:include: - local: gitlab/.gitlab-ci.yml variables: NANO_GITLAB_IMAGE: registry.gitlab.com/<group>/<project>/nano-analyzer:latest - Set CI/CD variables (masked):
NANO_GITLAB_API_TOKEN, a token that may write merge-request notes, and your LLM key, for exampleOPENROUTER_API_KEY. - Open a merge request. The job runs on merge-request pipelines only. It first updates the report with "This analysis is being updated...", then replaces it with a table of severity counts per file, the status, the time taken and the commit SHA.
How it behaves
- Scans only changed files with known source extensions (default mode
changed); onescan.pyprocess per file, run serially. - Skips a commit it has already analysed, and reopens the report discussion if someone resolved it.
- The comment shows counts only; the full findings are in the job artifacts under
.nano-analyzer/gitlab/(kept one week).
| Variable | Purpose |
|---|---|
NANO_GITLAB_SCAN_MODE | changed (default) or all |
NANO_GITLAB_SCAN_TARGET | What to scan in all mode (default .) |
NANO_GITLAB_CHANGED_BASE | Explicit base commit or branch for the diff |
NANO_GITLAB_ENFORCE_ONCE_PER_COMMIT | Skip commits already reported (default true) |
NANO_GITLAB_REOPEN_RESOLVED_THREAD | Reopen a resolved report discussion (default true) |
NANO_GITLAB_RESOURCE_GROUP | Lock key that keeps scans one-at-a-time |
NANO_GITLAB_MODEL, NANO_GITLAB_PARALLEL, NANO_GITLAB_TRIAGE_ROUNDS | Passed through to the scanner |
Security note: keep the API and LLM keys masked and protected, and do not expose them to pipelines from untrusted forks.
5. GitHub pull requests Draft
Our private codebase contains a GitHub Action definition that runs the scanner and keeps one comment up to date on the pull request. It is a draft: the inputs for target, model, parallelism, size limit, triage settings, confidence threshold and repo directory are wired to the scanner, but several others are declared and not wired yet: format, scope, provider, SARIF upload, artifact upload and the fail-* policy. It also runs scan.py from the caller's workspace, so it is not yet usable as a drop-in action from another repository.
Until that is tidied up, the simplest route is to run the command-line scanner in your own workflow. This is an untested sketch:
name: nano-analyzer
on: pull_request
permissions:
contents: read
jobs:
scan:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: actions/checkout@v4
with:
repository: Huge/nano-analyzer
path: .nano-analyzer
- name: Scan changed source files
env:
OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}
run: |
git diff --name-only --diff-filter=AMRT "origin/${{ github.base_ref }}...HEAD" \
| grep -E '\.(c|h|cc|cpp|go|py|js|ts|rs|java)$' > changed.txt || true
[ -s changed.txt ] || { echo "No source changes"; exit 0; }
xargs -a changed.txt python3 .nano-analyzer/scan.py \
--model qwen/qwen3-32b --repo-dir . --output-dir scan-results
- uses: actions/upload-artifact@v4
if: always()
with:
name: nano-analyzer-results
path: scan-results
Pull requests from forks do not receive repository secrets by default. That is a feature: do not work around it with pull_request_target plus a checkout of the pull request's code.
6. The planned @budgetscan command Planned
The idea: a reviewer comments on a pull or merge request and BudgetScan plans a scan that fits the cap and deadline. Nothing below is live yet; the syntax may change.
| Comment | Intent |
|---|---|
@budgetscan $0.25 | Quick pass over the changed files with a cap of $0.25 |
@budgetscan $0.50 --max-delay 5m | Deeper pass that stops starting new work as the 5-minute deadline nears and posts a partial report of what was covered |
@budgetscan $5.00 --depth deep | Spend a larger cap on wider context and more triage rounds |
- Amounts are in US dollars;
$0.50and.5mean the same. Deadlines look like90s,5mor1h. - The scan should never cost more than the cap. If the cap is too small for even one file, you get a clear "insufficient budget" report and nothing is charged.
- Failures are rare, and a failed scan is never charged. Full billing terms will be published before launch.
- Real per-scan costs will be published after we have measured them on benchmarks.
7. Where your code goes
The short version: hosted scans run on our partner Kosmik Compute (Prague, zero data retention), or you run the open-source scanner against your own endpoint on your own infrastructure. The details are in Where does my code go?
8. Limitations
- False positives. AI scanners flag things that turn out to be harmless. Multi-round triage helps, and we work hard on keeping false positives to a minimum, but always verify findings by hand.
- False negatives. Entire classes of problems (logic bugs, race conditions, cryptographic and authentication flaws) can be missed. A clean scan does not mean the code is safe.
- One file at a time. Each file is scanned on its own, so bugs that depend on interactions between files are likely to be missed.
- Model-dependent. Different models find different things and make different mistakes.
9. Questions and early access
Want the private integrations, have a question, or found something wrong in this guide? Email us (dev@ehlas.cz and safeAIwork@gmail.com).