308 Commits
Author SHA1 Message Date
Lum1104 9d1318a0e7 feat(i18n): add Russian language support 2026-05-19 15:01:43 +08:00
Lum1104 0e39f227a4 chore(release): bump version to 2.7.3
Ships the fingerprints baseline fix (e7af9ae): every install since
2.7.0 had a broken Phase 7 step 2.5 that threw TypeError on the first
/understand run and left fingerprints.json empty/missing, which made
every subsequent auto-update escalate to FULL_UPDATE. This release
replaces the LLM-written script with a bundled build-fingerprints.mjs
and reorders Phase 7 to write fingerprints before meta.json.

Anyone upgrading from 2.7.0–2.7.2 should re-run /understand --full
to regenerate a valid baseline.
2026-05-18 10:13:22 +08:00
Lum1104 e7af9ae35e fix(skills/understand): bundle build-fingerprints.mjs and reorder Phase 7
The Phase 7 step 2.5 code example in SKILL.md called
buildFingerprintStore() with 2 arguments, but the real signature
requires 4 (projectDir, filePaths, registry: PluginRegistry,
gitCommitHash: string). It also omitted the required
`await TreeSitterPlugin.init()`. Any LLM following the example
threw TypeError on `registry.analyzeFile()` and never produced a
baseline — which is why fingerprints.json never existed in a usable
form after a fresh /understand, and is the root cause behind
issue #152's "every auto-update escalates to FULL_UPDATE" cascade.

Replace the LLM-written script with a bundled `build-fingerprints.mjs`
that mirrors `extract-structure.mjs`: resolves @understand-anything/core
via createRequire, initializes TreeSitterPlugin + PluginRegistry
correctly, calls buildFingerprintStore with all four arguments, and
persists via saveFingerprints. Smoke-tested on this repo (3 files,
correct functions/classes/imports extracted).

Reorder Phase 7 so fingerprints are written BEFORE meta.json. If
fingerprint generation fails, the new step explicitly says to abort
Phase 7 — meta.json must not advance without a valid baseline, or
the next auto-update sees a fresh commit hash with no fingerprints
and classifies every file as STRUCTURAL.

Affects every install since 2.7.0 (when the broken example was
introduced). Users running /understand --full on 2.7.3+ will get
a usable fingerprints.json on the first try.
2026-05-18 10:13:02 +08:00
Lum1104 97fa2f3cab chore(release): bump version to 2.7.2
Ships two auto-update fixes:
- #153 (5304ff0): apply .understandignore in Phase 0 so user-excluded
  paths don't inflate the structural-change count.
- #152 (dd8b724): LOAD-PATCH-SAVE template for Phase 3d fingerprints
  merge, with guard against silent load failure.
2026-05-18 09:59:05 +08:00
Lum1104 dd8b724c99 fix(hooks/auto-update): make fingerprints merge unambiguous in Phase 3d
Fixes #152. Phase 3d step 3 instructed the LLM to "merge with existing
fingerprints (keep unchanged files as-is)" but the prose was vague
enough that the LLM-written script frequently wrote only the freshly
re-analyzed batch entries to fingerprints.json, discarding every other
file's fingerprint. The next auto-update saw N-batch_size files with
no stored fingerprint → classified as STRUCTURAL → exceeded the 30-file
threshold → FULL_UPDATE permanently, burning hundreds of thousands of
tokens on every subsequent commit.

Replace the four-bullet description with an explicit LOAD-PATCH-SAVE
script template:

  1. LOAD ALL existing entries from fingerprints.json (never skip).
  2. PATCH or REMOVE each path in filesToReanalyze (inline deletion
     handling so the spec doesn't need a separate deletedFiles list).
  3. GUARD: if the file existed and was non-empty but loaded as {},
     abort the write — silent load failure would otherwise clobber
     every fingerprint.
  4. SAVE the full dict back.

The reporter's dry-run showed this restores 81/97 files to COSMETIC
classification on their project (zero LLM tokens) instead of all 97
incorrectly forced into STRUCTURAL.

Note: a related ordering bug exists in skills/understand/SKILL.md
Phase 7 (meta.json written before fingerprints.json — silent failure
in step 2.5 leaves stale fingerprints). That's a separate fix in a
different file and is intentionally not bundled here.
2026-05-18 09:58:45 +08:00
Lum1104 5304ff06f3 fix(hooks/auto-update): apply .understandignore exclusions in Phase 0
Fixes #153. Phase 0 step 7 filters changed files to source extensions
only and never reads `.understandignore`, so files in user-excluded
paths (migrations, vendored code, tests) count as structural changes
and can spuriously escalate the action to FULL_UPDATE. The reporter
saw 50 → 38 structural files after applying their ignore patterns
(below the 30-file FULL_UPDATE threshold, ARCHITECTURE_UPDATE would
have sufficed).

Add step 9 that delegates to `createIgnoreFilter` from
`@understand-anything/core` via $CLAUDE_PLUGIN_ROOT. Same code path
as /understand's project-scanner Step 2.5, so the auto-update honors
the exact same patterns (hardcoded defaults + user .understandignore
files at both standard locations + `!` negation semantics).

If $CLAUDE_PLUGIN_ROOT can't be resolved, fail loud rather than
silently skipping — a silent skip reproduces the original bug.
2026-05-18 09:58:11 +08:00
Lum1104 2da74848e5 chore(release): bump version to 2.7.1
Ships the fixes that landed on main after the 2.7.0 cut:

- #139 (f3ea1a3): understand-knowledge — Windows path separators in
  wikilink resolution + omit empty `category` so KnowledgeMetaSchema's
  `z.string().optional()` no longer drops every article node. Closes #151.
- #147 (fafb888): understand-domain — resolve $PLUGIN_ROOT at runtime
  for symlink installs.
- f71bad5: understand — persist canonical edge direction during merge,
  fixing the 153k auto-correction cascade. Closes #140.
2026-05-18 09:48:44 +08:00
Lum1104 f71bad5267 fix(skills/understand): persist canonical edge direction during merge
Fixes #140. merge-batch-graphs.py already defaulted missing `direction`
to "forward" when building the dedup key, but the value was never written
back onto the edge — so the generated knowledge-graph.json shipped without
the field and the dashboard validator emitted one auto-correction per
edge (153k on the reporter's Go codebase).

Mirror the dashboard schema validator at packages/core/src/schema.ts:
lowercase the value, map "both"/"mutual" → "bidirectional", fall back to
"forward" for missing or invalid values, and persist the result onto the
edge before it enters edges_by_key. This also closes a latent dedup leak
where "Forward" and "forward" (or "both" and "bidirectional") would have
produced separate dedup keys.
2026-05-18 09:27:06 +08:00
Yuxiang LinandGitHub a5644ca3b0 Merge pull request #139 from nieao/fix/windows-compat
fix(understand-knowledge): Windows path + zod schema null compatibility
2026-05-17 22:07:38 +08:00
Yuxiang LinandGitHub cde871990c Merge pull request #147 from rustanacexd/fix/understand-domain-plugin-root
fix(skills/understand-domain): resolve plugin root for agent prompt loading (#146)
2026-05-13 10:20:33 +08:00
Rustan Corpuz fafb888422 fix(skills/understand-domain): resolve plugin root at runtime for symlink installs
Fixes #146. Ports the $PLUGIN_ROOT resolution pattern from understand/SKILL.md
to understand-domain/SKILL.md, including:
- Symlink resolution for ~/.agents/skills/understand-domain
- Copilot fallback for ~/.copilot/skills/understand-domain
- Detailed error diagnostics listing all checked paths
- Phase 4 agent prompt path now uses $PLUGIN_ROOT/agents/domain-analyzer.md
2026-05-12 15:00:55 +08:00
Lum1104andClaude Opus 4.7 04c84ab4a4 chore(release): bump version to 2.7.0
Includes since 2.6.3: Hermes (#91), Cline (#116), KIMI CLI (#134)
platform support; dashboard ACCESS_TOKEN env override; README cleanup
(slogan rewrite, drop outdated overview gifs, move thanks to footer).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 14:27:12 +08:00
Yuxiang LinandGitHub 4c6f7c3e0b Merge pull request #142 from zhushen12580/feature/language-parameter
feat: Add --language parameter for localized content generation
2026-05-12 11:07:14 +08:00
zhushen 2083342199 fix: Address code review feedback for PR #142
- Remove invalid allowBuilds from pnpm-workspace.yaml (use onlyBuiltDependencies in root package.json)
- Use data-testid for search input selector (fixes / keyboard shortcut for non-English locales)
2026-05-12 11:02:13 +08:00
zhushen a3ec91bf39 feat(dashboard): Complete i18n translation for all UI components 2026-05-12 01:53:58 +08:00
Lum1104andClaude Opus 4.7 1779cbd9a7 feat(dashboard): allow ACCESS_TOKEN override via env var
Honor UNDERSTAND_ACCESS_TOKEN if set, falling back to the random 16-byte
hex token. Lets the dev token survive across server restarts so shared
dashboard URLs don't rot, and makes the auth path easier to script in
tests.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 22:02:58 +08:00
zhushen e1650f627c fix: Wrap MobileLayout with I18nProvider; use outputLanguage key in config
P1: MobileLayout was missing I18nProvider wrapper, causing useI18n
    to throw error on mobile devices. Now both desktop and mobile
    layouts are wrapped with I18nProvider.

P2: SKILL.md used 'language' key but Dashboard reads 'outputLanguage'.
    Fixed config.json key name to match ProjectConfig type definition.

All tests passed:
- Core: 670 tests
- Dashboard: 42 tests
2026-05-11 20:30:50 +08:00
zhushen 752fe59e0c feat(dashboard): Add i18n support for localized UI text
- Add outputLanguage field to ProjectConfig type
- Create /config.json endpoint in vite.config.ts
- Build locale files for 5 languages (en, zh, zh-TW, ja, ko)
- Add I18nProvider context and useI18n hook
- Update 5 components (ProjectOverview, NodeInfo, FileExplorer, FilterPanel, PersonaSelector)
- Dashboard reads language from config.json and displays localized UI

All tests passed:
- Core: 670 tests
- Dashboard: 42 tests
2026-05-11 19:00:05 +08:00
zhushen 656289121c feat: Add --language parameter for localized content generation
Adds --language parameter to /understand command to generate knowledge
graph content in user-specified language.

Changes:
- Update argument-hint and Options documentation in SKILL.md
- Add language parsing logic in Phase 0 (language normalization,
  config persistence, LANGUAGE_DIRECTIVE template)
- Inject language directive into agent dispatch prompts for all
  content-generating phases (Phase 1-5)
- Add language directive handling instructions in agent definitions
- Create locales/ directory with template files for:
  - English (en.md) - default
  - Chinese Simplified (zh.md)
  - Chinese Traditional (zh-TW.md)
  - Japanese (ja.md)
  - Korean (ko.md)

Locale files provide language-specific guidance for:
- Tag naming conventions
- Summary writing style
- Technical term handling
- Layer name translations

Closes #141
2026-05-11 12:20:05 +08:00
Yuxiang LinandGitHub 40519ee5ff Merge pull request #138 from voidborne-d/fix/worktree-paths
fix(skills): redirect PROJECT_ROOT out of git worktrees (closes #133)
2026-05-10 20:13:28 +08:00
Xingkai98 e29f461574 fix(dashboard): reset fn toggle on view switch; hide detail toolbar in domain view
- Issue 1: setDetailLevel now resets showFunctionsInClassView so the fn
  toggle doesn't resurrect when re-entering class view after a file-view
  round-trip.
- Issue 2: detail-level toolbar (Files/+Classes/fn) now gated on
  viewMode !== "domain" so it doesn't render in domain view where it has
  no effect.
2026-05-10 16:24:26 +08:00
Kai XingandGitHub f23291f9ea Merge branch 'Lum1104:main' into feat/dashboard-file-class-views 2026-05-10 15:40:06 +08:00
Yuxiang LinandGitHub a381c41ef6 Merge pull request #124 from tipich/fix/windows-pnpm10-compat
fix(skill): make /understand work on Windows + pnpm 10
2026-05-10 09:38:18 +08:00
nieaoandClaude Opus 4.7 f3ea1a3088 fix(understand-knowledge): Windows path + zod schema null compatibility
Four single-line fixes that make `/understand-knowledge` work end-to-end on Windows.

## Root causes

1. Path separator mismatch (3 occurrences in parse-knowledge-base.py)
   - `str(rel.with_suffix(""))` returns backslash-separated stems on Windows
     (e.g. `entities\foo`), while wikilinks always use forward slashes
     (`[[entities/foo]]`).
   - Result: name_map keys and article_ids hold `entities\foo`, while
     `resolve_wikilink()` looks up `entities/foo` -> 100% miss.
   - Tested on Windows 11 + Python 3.14: 151/151 wikilinks unresolved,
     0 edges built from wikilinks.

   Fix: use `rel.with_suffix("").as_posix()` in all three places
   (lines 235, 316, 330 on main).

2. Null vs missing field (1 occurrence)
   - `"category": category or None` writes `null` when category is empty.
   - `KnowledgeMetaSchema.category` in packages/core/src/schema.ts is
     `z.string().optional()`, which accepts `undefined`/missing but
     rejects `null`.
   - Result: every article node fails GraphNodeSchema validation in the
     dashboard (`Invalid input: expected string, received null`),
     all nodes get dropped, dashboard renders empty.

   Fix: omit the field when empty using dict spread.

## After

Tested with a 27-article Karpathy wiki on Windows:
- before: 151 unresolved wikilinks, 0 edges, 0 nodes rendered in dashboard
- after:  0 unresolved, 110 wikilink edges + 23 LLM-implicit edges,
          all 53 nodes (article/topic/entity/claim) render correctly

No behavior change on macOS/Linux: `as_posix()` is a no-op when the OS
already uses `/`, and dict spread produces the same key as the previous
truthy branch.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 15:38:13 +08:00
d 🔹 b962a3dc33 fix(skills): redirect PROJECT_ROOT out of git worktrees (#133)
When /understand or /understand-domain runs from a CWD inside an
ephemeral git worktree (the default for parallel-agent / isolation
sessions), every output file goes to the worktree path. Claude Code
deletes the worktree on session end, taking knowledge-graph.json,
domain-graph.json, meta.json, intermediate batches and ~hundreds of K
of analysis tokens with it.

Resolve PROJECT_ROOT through a worktree check before any output:
compare git rev-parse --git-dir against --git-common-dir; in a normal
checkout (and in a submodule) they're the same path, in a worktree they
differ and parent(--git-common-dir) is the main repo root.
UNDERSTAND_NO_WORKTREE_REDIRECT=1 opts out for the rare per-worktree
case.

- skills/understand/SKILL.md Phase 0 step 1: add the redirect after
  PROJECT_ROOT is set from $ARGUMENTS or CWD, so an explicit arg path
  is also rescued from a worktree but can be opted out of.
- skills/understand-domain/SKILL.md: add an explicit Phase 0 (it
  previously inferred "current project" implicitly), then thread
  $PROJECT_ROOT through Phases 2-5 so subsequent steps honor the
  redirect.
- New worktree-redirect.test.mjs: 5 vitest cases covering main repo,
  worktree root, worktree subdir, opt-out env var, and non-git CWD.
  Mirrors the bash snippet inline (no shared lib in this repo).

Submodule false-positive ruled out by probe — submodules see git-dir
== git-common-dir (both point at <super>/.git/modules/<name>).
2026-05-09 12:58:46 +08:00
Yuxiang LinandGitHub 3eb7700a8f Merge pull request #122 from Lum1104/feat/issue-113-tested-by-coverage
feat: deterministic tested_by edges + dashboard badge (#113)
2026-05-09 11:18:27 +08:00
Lum1104andClaude Opus 4.7 a4bdc1c99d fix(merge): keep max-weight tested_by edge in Pass 1 dedup (#113)
Codex P2: link_tests Pass 1 dropped duplicate (production, test) pairs
purely by arrival order — when two batches both emitted a tested_by
edge for the same pair with different confidences (0.3 vs 0.9), the
edge that happened to iterate first won. The general Step 6 deduper
at line 762 mirrors `weight > existing.weight` semantics but it only
ever saw one of the duplicates, so it couldn't rescue the heavier one.

Refactor Pass 1 to mirror Step 6's weight comparison locally:

  - Track `pair_to_idx` mapping each kept (prod, test) pair to its
    slot in the compacted edges list. On a duplicate, look up the
    existing kept edge and compare weights; if the new edge is
    strictly heavier, swap (if needed) and replace the slot. Tie or
    lighter → drop the new edge.
  - Defer the swap operation until we know an edge will survive — no
    point canonicalizing a doomed duplicate.
  - Track surviving swap pairs in a separate `swapped_pairs` set so
    the `swapped` counter reflects the FINAL output, not the wasted
    work on edges that were later replaced. This means: replacing a
    swapped edge with a heavier canonical one drops the swap from
    the count; replacing a canonical edge with a heavier swapped one
    adds it.
  - Extract the swap-in-place mutation into `_swap_tested_by_in_place`
    so it can be invoked from both code paths.

Five new unit tests cover all four weight-vs-direction combinations
plus a tie case (existing test_drops_duplicate_canonical_edges, which
still passes — tie → keep first, no swap counted).

microservices-demo regression check unchanged: 7 → 7 edges, 3 swapped,
0 dropped, 7 tagged.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 11:05:43 +08:00
Lum1104andClaude Opus 4.7 4bb22fd9af feat(merge): swap-then-supplement tested_by linker (#113)
Strip-and-rederive (current PR behaviour) drops real coverage signal on
projects whose test layout doesn't match a naming convention. On the
Google microservices-demo the LLM had emitted 7 valid tested_by edges
(3 with inverted direction); the strip pass dropped them and the path-
convention rederive could only re-pair 4 of them. Net: 7 → 4 edges,
3 production files lost their tested signal.

Replace strip-and-rederive with two-pass swap-then-supplement:

  Pass 1 — walk LLM tested_by edges. Canonical (production → test)
  edges pass through unchanged. Inverted (test → production) edges are
  flipped in place; description gets a `[direction corrected]` audit
  marker. Edges with no recoverable meaning (test↔test, prod↔prod,
  orphan endpoint, duplicate pair) are dropped.

  Pass 2 — for tests not yet paired by Pass 1, walk path-convention
  candidates and emit a fresh production → test edge for the first
  match. Pairs already covered by Pass 1 are skipped.

Tagging is consolidated into a final pass over all canonical edges so
production nodes get the "tested" tag whether the edge came from
Pass 1 (canonical / swapped) or Pass 2 (supplement).

Multi-language audit of production_candidates revealed three real-world
gaps surfaced by re-checking microservices-demo and common project
layouts:

  - JS/TS walk-out only handled `__tests__/`. Extended to also walk out
    of `<dir>/test/`, `<dir>/spec/`, and `<dir>/tests/` (some JS/TS
    projects use these instead of __tests__/).
  - Python walk-out only handled top-level `tests/`. Added in-package
    `<pkg>/tests/test_<name>.py` → `<pkg>/<name>.py` (Django app style
    and any project that colocates tests with the package).
  - C# only had sibling fallback. Added two new mirrors:
      * `<svc>/tests/X.cs` ↔ `<svc>/X.cs` and `<svc>/src/.../X.cs`
        (microservices-demo cartservice exact layout).
      * `<App>.Tests/Foo/BarTests.cs` ↔ `<App>/Foo/Bar.cs`
        (.NET sibling-project convention).

Go is intentionally not changed — the "one _test.go covers several
.go files in the same package" pattern is now solved by Pass 1
(swapping LLM edges), not by trying to invent multi-pair path heuristics.

The file-analyzer prompt is updated: the `tested_by` row is restored
in the schema table because we now use those edges as evidence (Pass 1
canonicalizes the direction). The note explains direction will be
auto-corrected so the LLM doesn't need to be defensive about it.

link_tests now returns a 4-tuple (added, dropped, tagged, swapped);
the merge_and_normalize report distinguishes "edges produced
(supplement)" from "edges flipped" from "edges dropped".

Real-world validation on microservices-demo:
  before:   7 tested_by edges, 3 inverted, 0 tagged
  after PR: 4 tested_by edges, 0 inverted, 4 tagged   ← strip-and-rederive
  this:     7 tested_by edges, 0 inverted, 7 tagged   ← swap-then-supplement

Tests: 47 pass (was 37). New cases cover all swap branches, the
shippingservice "one test, many sources" regression, and each new
language pattern.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 10:45:22 +08:00
Xingkai98andClaude Opus 4.6 c1bc4bfac1 feat(dashboard): add file/class dual-view toggle to reduce graph clutter
Add detailLevel state ("file" | "class") to separate architecture-level
file dependencies from code-structure class views. File view shows only
file nodes and file→file edges (imports/depends_on), eliminating ~85%
of nodes/edges that previously caused severe zoom/pan lag. Class view
adds class nodes via contains edges with an optional function toggle.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-09 01:57:10 +08:00
Yuxiang LinandGitHub 5a2ffb1e0f feat: customizable heading font via theme settings (#121)
Add a `headingFont` option to ThemeConfig that lets users switch
heading typography between Serif (default), Sans, and Mono via the
existing Theme Picker UI.

- New `--font-heading` CSS custom property (defaults to `--font-serif`)
- Theme engine applies the selected font on config change
- ThemePicker gets a "Heading Font" toggle section
- All 15 component files updated from `font-serif` to `font-heading`
- Selection persists in localStorage alongside other theme settings
- Backwards-compatible: existing configs without `headingFont` default to serif

Closes #120
2026-05-08 22:20:47 +08:00
Lum1104andClaude Opus 4.7 05fd42343c fix(merge): recover imports edges file-analyzer batches drop
A controlled-experiment audit on a 1240-file Python project (opensre)
showed that 27.2% of resolved-internal imports never made it from
project-scanner's `importMap` into the final knowledge graph. Of the
404 source files with internal imports, 91 ended up with ZERO imports
edges in the graph despite their `file:` node being present (consistent
with main-session orchestrator dropping the entry from `batchImportData`
during batch construction), and 104 had partial coverage (consistent
with file-analyzer agent dropping rows during edge enumeration).
GitHub issue #128 reported the same failure mode at 16-21% on a Go
monorepo.

The fix has two layers:

1. `merge-batch-graphs.py` now runs a deterministic recovery pass
   after merge: for every `(source, target)` in scan-result.json's
   `importMap` whose source `file:` node exists in the assembled graph
   and whose target `file:` node also exists, emit an `imports` edge
   if the batches didn't already. Recovered edges are tagged
   `recoveredFromImportMap: true` so downstream consumers can audit
   which edges came from the deterministic source vs. agent emission.
   The merge report logs the recovered count plus how many importMap
   entries were skipped because their source/target had no graph node.

2. `file-analyzer.md` rewrites the imports edge rule to demand 1:1
   emission with a self-check: "the number of `imports` edges in your
   output MUST equal `sum(batchImportData[file].length)` across the
   batch's code files". This drives the agent to enumerate every row
   instead of summarizing — recovery should report 0 when this works.

Tests: +6 cases covering the recovery path — drops, no-double-emit,
missing source/target nodes, missing scan-result.json (incremental
update), and self-import suppression. 770 passing (was 764).

Bumps version to 2.6.3 across the five tracked manifests.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-08 19:35:34 +08:00
tipichandClaude Opus 4.7 366368face fix(skill): make /understand work on Windows + pnpm 10
Two bugs blocked /understand from running properly:

1. extract-structure.mjs:34,37 — `await import()` was called with raw
   absolute paths returned by `require.resolve()` and `path.resolve()`.
   On Windows those start with "C:\..." which Node 24's ESM loader
   parses as URL scheme "C:" and rejects with
   ERR_UNSUPPORTED_ESM_URL_SCHEME. Wrap both with `pathToFileURL().href`
   so the loader receives a proper file:// URL.

2. tree-sitter native build scripts were skipped at install time
   because pnpm 10 blocks postinstall scripts by default. Added the
   parser packages plus esbuild and sharp to `pnpm.onlyBuiltDependencies`
   so they compile during `pnpm install` and `analyzeFile` / fingerprint
   generation actually work.

Without these, file-analyzer subagents fall back to direct file reads
and the importMap stays empty across all batches, which forces
assemble-reviewer to grep-recover hundreds of cross-batch edges by
hand. They also block the incremental update path entirely because
fingerprint generation crashes inside `buildFingerprintStore`.

Verified end-to-end on a 1190-file Laravel codebase: tree-sitter PHP
parses Booking.php into 18 typed functions, fingerprint baseline
builds in 1.56s, change detection correctly classifies cosmetic
vs structural diffs.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-08 13:48:36 +03:00
Lum1104andClaude Opus 4.7 c49c46d974 fix(pipeline): close 12 sources of silent data loss in graph extraction
A deep audit of the project-scanner → file-analyzer → merge pipeline
turned up a wide range of silent data-loss bugs. Each one alone is
small; together they were producing graphs with very few import edges,
missing sub-file nodes for non-code formats, and inconsistent metrics.

Root-cause fixes (high impact):

- project-scanner.md: extend import-pattern table to resolve absolute
  imports for Python (`from a.b.c import x`), TS/JS (tsconfig.json
  paths/baseUrl aliases), Java/Kotlin (`com.foo.Bar` ↔ file paths),
  Ruby (`require 'foo/bar'` load-path), PHP (composer PSR-4 namespaces),
  and C/C++ (`#include` headers). Was relative-only, which produced
  empty importMap entries for the majority of real projects.
- project-scanner.md: add `.ps1`, `.bat`, `.cmd`, `.jsonc` to language
  table; require non-null `language` field with an explicit fallback.
- file-analyzer.md: document `sections`, `definitions`, `services`,
  `endpoints`, `steps`, `resources` in the extraction-output schema and
  spell out the sub-file node-creation rules per category. Was missing,
  so per-table / endpoint / resource nodes were never created from
  SQL / OpenAPI / Terraform / K8s / Dockerfile parser output.
- file-analyzer.md: add explicit source-reading fallback rules for
  PowerShell, Batch, Bash, Swift, Kotlin (no tree-sitter coverage).
- yaml-parser: declare `kubernetes`, `docker-compose`, `github-actions`,
  `openapi` languages so files the language-registry tags with those
  ids actually get section extraction. Recognize quoted top-level keys
  (e.g. `"on":` in GitHub Actions). Emit one section per entry for
  array-root YAML documents.
- json-parser: declare `json-schema`, `openapi`; add `stripJsoncSyntax`
  helper that removes line / block comments and trailing commas before
  parse so `.jsonc` files (wrangler, tsconfig with comments) parse cleanly.
- shell-parser: declare `jenkinsfile`. Tighten function-detection regex
  to require a reachable `{` brace so `name() echo hi` and patterns
  appearing inside heredocs are no longer false-positives.
- markdown-parser: track fenced-code-block state and skip headings
  inside ``` / ~~~ blocks (`# install` shell comments were being
  emitted as level-1 sections).
- merge-batch-graphs.py: add `article`, `entity`, `topic`, `claim`,
  `source` to VALID_NODE_PREFIXES and TYPE_TO_PREFIX so knowledge-base
  node types stop being flagged unknown / coerced to `file:`. Add
  `direction` to the edge dedup key so `forward` and `bidirectional`
  variants of the same (src, tgt, type) don't overwrite each other.
  Use a placeholder in bare-id fallback when `filePath` is missing on
  function/class nodes so unrelated `parse()` functions don't merge.
- typescript-extractor: actually compute `isDefault` for default
  exports (was always emitted as `false` from buildResult).
- extract-structure.mjs: match `wc -l` semantics for `totalLines` so
  the scanner's `sizeLines` and the extractor's `totalLines` agree on
  POSIX text files. Filter the parser-imports fallback to relative-only
  so `importCount` semantics stay *internal-import* whether the scanner
  resolved them or not. Drop unused `isCode` local.

Tests: +19 cases covering JSONC parsing, markdown fenced-code skip,
YAML quoted-keys / array-root, shell function false-positives,
extract-structure import fallback semantics + totalLines off-by-one.
764 passing (was 745).

Bumps version to 2.6.2 across the five tracked manifests.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 21:44:24 +08:00
Lum1104andClaude Opus 4.7 7834f9bc75 fix(file-analyzer): preserve language and fall back when imports unresolved
Two bugs surfaced when analyzing Python projects that use absolute imports:

- The dispatch prompt (SKILL.md) and file-analyzer agent omitted the
  per-file `language` field, so `extract-structure.mjs` received null and
  passed it through to the graph.
- `extract-structure.mjs` used `if (importPaths)` to decide whether to
  trust pre-resolved imports. Empty arrays are truthy, so files where the
  project scanner could not resolve any imports (e.g. Python absolute
  imports) clobbered the parser's import count with 0, never falling
  back to tree-sitter's own analysis.

Bumps plugin version to 2.6.1 across the five tracked manifests and adds
unit tests for `buildResult` covering language pass-through and the
importCount fallback paths. To make the script testable, `buildResult` is
now exported and the CLI invocation is guarded so importing the module
no longer triggers `main()`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 21:02:47 +08:00
Lum1104andClaude Opus 4.7 fdfa331009 fix(merge): coerce malformed tags before adding "tested" (#113)
Codex flagged that prod_node.setdefault("tags", []) returns the existing
value when the key is present, so a raw LLM batch with tags=None or
tags="some string" would crash the whole merge on the next "tested" not
in tags membership check.

The TypeScript autoFixGraph normalizer that handles this case runs
downstream of merge-batch-graphs.py, not before it, so the Python side
has to defend itself. Coerce non-list tags to a fresh [] before the
membership/append.

Regression test exercises None / comma-string / single-string / int /
dict inputs.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 10:25:56 +08:00
Denis BalanandGitHub 53248fea8b Merge pull request #117 from DenisBalan/patch-1
Fix dashboard URL format in vite.config.ts
2026-05-07 10:14:13 +08:00
Lum1104andClaude Opus 4.7 6c257a55f0 feat(dashboard): tested badge on node cards (#113)
Render a small green dot next to the complexity badge whenever a node's
tags contain "tested" — surfacing the deterministic linker's signal so
users can see at a glance which files have paired tests.

Plumb node.tags through both CustomNodeData construction sites in
GraphView.tsx; KnowledgeGraphView.tsx already passes tags.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 22:20:40 +08:00
Lum1104andClaude Opus 4.7 ba9eeab5bd refactor(merge): polish tested_by linker per code review (#113)
- is_test_path: collapse 7 per-language conditional blocks into a
  data-driven _TEST_NAME_PATTERNS table; JS/TS infix stays inline
- production_candidates: extract _join + module-level _add_unique to
  drop the nested closure and the repeated trailing-slash idiom
- Drop dead _TEST_DIR_SEGMENTS constant and the local _splitext
  reimplementation; use os.path.splitext
- link_tests: drop the impossible-malformed-tags guard, tighten the
  docstring, change edge description to "Path-based pairing
  (deterministic)", drop redundant break comment
- Trim Step 5b inline block that duplicated the module-level header
- Convert file-analyzer Note from blockquote to bold paragraph to
  match surrounding prompt style

Tests: split the strip-edges test from the unrelated-edges-survive
test, add empty-input and missing-filePath cases, pin sibling-before-
walkup and sibling-before-mirror priority order, drop brittle report
text assertion. 36 tests, all passing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 22:13:37 +08:00
Buzzwoo Team E-Com 92a13eadb5 feat: customizable heading font via theme settings
Add a `headingFont` option to ThemeConfig that lets users switch
heading typography between Serif (default), Sans, and Mono via the
existing Theme Picker UI.

- New `--font-heading` CSS custom property (defaults to `--font-serif`)
- Theme engine applies the selected font on config change
- ThemePicker gets a "Heading Font" toggle section
- All 15 component files updated from `font-serif` to `font-heading`
- Selection persists in localStorage alongside other theme settings
- Backwards-compatible: existing configs without `headingFont` default to serif

Closes #120
2026-05-06 16:13:28 +02:00
Lum1104andClaude Opus 4.7 f661a03376 docs(agents): drop tested_by from file-analyzer prompt (#113)
Now that the merge script produces tested_by edges deterministically
from path conventions, the LLM should not emit them — its direction is
unreliable across batches and any emitted edges are stripped on merge.

- Remove tested_by row from file-analyzer's edge table.
- Add a note pointing to the deterministic linker.
- Document the new behaviour in the merge section of SKILL.md.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 22:00:41 +08:00
Lum1104andClaude Opus 4.7 d9aa9cbc89 feat(merge): deterministic tested_by linker (#113)
The file-analyzer LLM only sees the production↔test relationship when
analyzing a test file (production files don't import their tests), so
its emitted direction was unreliable across batches and recall was
massively undercounted (~7% on a real Nuxt 4 + Directus repo).

Move tested_by production entirely into the merge step. The linker:

- Strips every tested_by edge from batch input (LLM direction unreliable).
- Indexes file:* nodes and classifies each path as test or production.
- For each test, walks ordered candidate production paths (sibling
  de-infix, __tests__/ walk-out, mirrored tests/→{src,app,lib,<root>}
  tree, Maven/Gradle src/test/...→src/main/...).
- Emits canonical production → test edges and tags production nodes
  "tested".

Supported conventions: JS/TS family (.test/.spec), Go (_test.go),
Python (test_*.py, *_test.py), Java (*Test/*Tests/*IT.java), Kotlin
(*Test/*Tests.kt), C# (*Test/*Tests.cs), C/C++ (test_*, *_test).

Stdlib only, type-hinted in existing style. Hooked into
merge_and_normalize between node dedup (Step 5) and edge dedup
(Step 6). Reports drops under "Fixed" and additions under a new
"Tested-by linker" section.

Tests cover path classification, candidate generation, full link_tests
behaviour (forward direction, idempotence, LLM-edge stripping,
test-to-test rejection), and the merge integration. 31 cases, stdlib
unittest, runnable with `python -m unittest test_merge_batch_graphs.py`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 22:00:34 +08:00
Lum1104andClaude Opus 4.7 880e223bb4 feat(dashboard): add mobile layout and responsive fixes
Bumps to 2.6.0.

- MobileLayout activates via useIsMobile at <768px with bottom-tab
  navigation (Graph/Info/Files); panes stay mounted (visibility
  toggle) to preserve ReactFlow dimensions and FileExplorer state.
- MobileDrawer holds persona, view mode, diff, node-type filters,
  layers, and tool buttons (Filter/Export/Path/Theme/Help).
- Selecting a node auto-pivots to Info; CodeViewer is always
  fullscreen on mobile; SearchBar collapses to a 🔍 toggle.
- Homepage Hero/Footer/Install responsive: drop nowrap on title and
  tagline, stack title spans for editorial wrap, full-width CTAs at
  <480px, narrow-width spacing refinements.
- Desktop dashboard: sidebar telescopes 260/300/360px, header gaps
  tighten, Path button label collapses to icon at narrow widths.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 16:25:56 +08:00
Lum1104andClaude Opus 4.7 6e0d1c11ea feat(dashboard): add favicon to match homepage branding
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 19:52:08 +08:00
Lum1104andClaude Opus 4.7 61356da3dc fix(dashboard): suppress TourFitView overlay flicker after fallback
Follow-up on the Codex P2: now that `useNodes` is in the effect deps,
every node update during a step that already timed out re-enters the
poll, sets `tourFitPending=true`, runs RAF for 4s, hits the silent
fallback path, and clears the flag. Visually the "Locating tour
highlight…" overlay would flash on every reflow even though the user
has already given up waiting. Skip the pending flag once
`fallbackKeyRef` matches the current step — the retry still runs
silently so a late Stage 2 can still upgrade to the proper fit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 19:27:30 +08:00
Lum1104andClaude Opus 4.7 4e1b83a35d chore: bump to 2.5.1
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 19:24:15 +08:00
Lum1104andClaude Opus 4.7 75368dc103 fix(dashboard): TourFitView timeout no longer freezes refit
Codex review on PR #114: when the RAF poll window expires before
highlighted nodes have been measured, the timeout fallback was setting
`fittedKeyRef.current = targetKey`, marking the step as fitted even
though the proper highlight fit never ran. If Stage 2 layout landed
after the 4s cap, the effect early-returned on the next nodes update
because the target key already matched, so the camera stayed pinned to
the fallback layer fit instead of zooming onto the actual highlights.

Fix:

  - Subscribe to React Flow's user-node array via `useNodes()` so the
    effect re-fires when Stage 2 finally produces the highlighted ids
    after the per-step RAF poll has already given up.
  - On timeout, pan into the layer for usability but do NOT set
    `fittedKeyRef`. The next nodes update gets another shot at the
    highlight fit, and on success `fittedKeyRef` records the proper fit.
  - Use a separate `fallbackKeyRef` to ensure the fallback `fitView`
    fires at most once per step — without this, every subsequent nodes
    update during the unready window would trigger a viewport jump.
  - Reset both refs when `tourHighlightedNodeIds` clears (stop tour).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 19:24:09 +08:00
Lum1104andClaude Opus 4.7 f58f0edd62 fix(dashboard): clear pendingFocusContainer on layer/state resets
Codex review on PR #114 flagged that `layerResetIfChanged` cleared
`containerLayoutCache` and `expandedContainers` but left
`pendingFocusContainer` intact. Because container ids collide across
layers (the very reason the cache reset exists), a manual expand in
layer A that hadn't yet hit its 1.2s clear timer could leak its id
into layer B's namespace and recenter the viewport on an unrelated
container right after navigation.

The same hazard applies to every other reset path that drops the
container caches. Add `pendingFocusContainer: null` to all of them:

  - layerResetIfChanged (tour cross-layer reset, the originally flagged
    site)
  - drillIntoLayer
  - navigateToOverview
  - setFocusNode
  - setPersona
  - setGraph
  - toggleNodeTypeFilter
  - clearContainerLayouts

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 19:21:06 +08:00
Lum1104andClaude Opus 4.7 4b86c696a5 fix(dashboard): tour navigation glitches across layers
Four related issues that surfaced while walking the Learn-mode tour
through a multi-layer project (microservices-demo):

1. Tour auto-expand never released. The tour effect that expands
   highlighted nodes' containers had no corresponding collapse when the
   step changed, so containers accumulated open as the user advanced.
   Track the set of containers we expanded and release any not needed
   by the current step; user-toggled containers are never tracked here,
   so they're never auto-collapsed.

2. Manual container toggle yanked off-screen. Stage 2 reflow shifted
   the just-clicked container away from the cursor. `toggleContainer`
   now records `pendingFocusContainer` on expand; GraphView locks the
   viewport onto that container's centre with the current zoom so it
   appears to expand in place.

3. Tour fitView fired before highlighted children existed. A single
   RAF after `tourHighlightedNodeIds` change wasn't enough — child
   nodes only appear once Stage 2 layout writes
   `containerLayoutCache`, and React Flow only knows their absolute
   position after a measure pass. `useNodes()` doesn't fire on
   measure completion, so we poll `getInternalNode().measured` each
   frame (up to ~4s) and call `fitView({ nodes })` once every
   highlight is measured, with `maxZoom: 1.2 / minZoom: 0.4`. While
   waiting, a new `tourFitPending` flag drives a "Locating tour
   highlight…" overlay so the user knows the layout is still settling.

4. Cross-layer tour transitions reused stale Stage 2 cache. Container
   ids derive from per-layer state (folder names in folder strategy,
   `container:cluster-N` in community strategy) and collide across
   layers — API Contracts and Load Testing both produce
   `container:cluster-0`. `setTourStep` / `nextTourStep` /
   `prevTourStep` / `startTour` didn't reset the container caches the
   way `drillIntoLayer` does, so when tour crossed a layer the new
   layer's expanded containers hit the previous layer's cache, Stage 2
   skipped its rerun, and children never showed. Extracted
   `layerResetIfChanged` and applied it in all four tour actions.

Verified end-to-end against microservices-demo via headless Chrome:
all 15 tour steps now expand the right container(s), zoom onto the
referenced files, and collapse the previous step's auto-expansions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 19:06:39 +08:00
d 🔹andClaude Opus 4.7 988533a550 fix(dashboard): preserve any-layer-wins membership for filterNodes
Reviewer @Lum1104 (PR #112) caught a silent semantic regression: the new
`filterNodes` reads layer membership through `nodeIdToLayerId.get(node.id)`,
which is first-wins. The pre-#112 path was any-layer-wins —
`layers.some(layer => filters.layerIds.has(layer.id) && layer.nodeIds.includes(node.id))`.
For a node X listed in both L1 and L2 with only L2 selected, the old code
kept X; the new code dropped it. The schema permits multi-layer membership,
so this was a behavior change, not a bug fix.

Fix: keep two distinct indexes in the store. Both are rebuilt once on
`setGraph`, so the O(1)-per-node performance win from #112 is preserved.

  - `nodeIdToLayerId: Map<string, string>` — first-matching-layer wins.
    Drives navigation (drillIntoLayer / tour step → layer / sidebar
    history) where one canonical layer is the right answer. Unchanged.

  - `nodeIdToLayerIds: Map<string, Set<string>>` — every layer the node
    belongs to. Drives `filterNodes` membership checks. Restores
    any-layer-wins exactly.

`filterNodes` now iterates the (small) layer-id set per node looking for
intersection with `filters.layerIds`. ExportMenu reads
`nodeIdToLayerIds` from the store.

Verified locally:

  - Added `filters.test.ts` regression: node in (L1, L2) with only L2
    selected must pass. Failed against the first-wins implementation;
    passes now.
  - `pnpm --filter @understand-anything/dashboard test` — 42 / 42 pass
    (was 41; +1 multi-layer regression test; perf-guard at 100 layers ×
    100 nodes still <50 ms).
  - `pnpm --filter @understand-anything/dashboard exec tsc --noEmit` — clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 17:16:42 +08:00
d 🔹andClaude Opus 4.7 44e1fee31b fix(dashboard): O(N+K) per-layer aggregations, kill quadratic Array.includes (#102)
Three hot paths in the dashboard ran `layer.nodeIds.includes(node.id)`,
which is O(K) per check. Combined with their enclosing loops they
collectively spent quadratic time per render of the overview / per
filter recompute / per node-selection event. On the 4.8 MB knowledge
graph reported in #102, the overview render alone took ~470 ms of
synchronous main-thread work before ELK / React Flow ran — long
enough for the page to register as unresponsive.

Fix: precompute two indexes once when a graph is loaded.

  - `nodesById: Map<string, GraphNode>`
  - `nodeIdToLayerId: Map<string, string>`  (first layer wins, matching
    prior `findNodeLayer` semantics)

Both live in `useDashboardStore` and are rebuilt by `setGraph`. The
three call sites:

1. `useOverviewGraph` (GraphView.tsx) — per-layer complexity aggregation
   moved into a new `computeLayerStats(layer, nodesById)` helper that
   walks `layer.nodeIds` instead of filtering all `graph.nodes`. Search
   match counts now read straight from `nodeIdToLayerId` instead of
   rebuilding a layer index on every searchResults change.

2. `filterNodes` (utils/filters.ts) — takes `nodeIdToLayerId` instead of
   `Layer[]`; the layer-membership check is one Map.get() per node.
   Updated `ExportMenu.tsx` caller to pass the store-level index.

3. `findNodeLayer` (store.ts) — replaced with `nodeIdToLayerId.get()` at
   the four call sites. `navigateTourToLayer` helper updated to take the
   index rather than the whole graph.

Behavior is preserved exactly:

  - "First layer wins" semantics for nodes that appear in multiple
    layers (#102 schema doesn't forbid this).
  - 30 % aggregate-complexity threshold pinned by tests.
  - Layer filter that excludes layer-less orphans, but ungated when
    no layers are selected.

Verified locally:

  Bench (`scripts/benchmark-aggregations.mjs`, node 22):
    100 layers × 200 nodes (#102 shape):  475 ms → 2 ms  (232× faster)
    50  layers × 200 nodes:               116 ms → 0.6 ms (190× faster)
    30  layers × 100 nodes:                12 ms → 0.2 ms (63×  faster)

  Tests: `pnpm --filter @understand-anything/dashboard test`
    24 → 41 pass (+17 new tests across `layerStats.test.ts` and
    `filters.test.ts`, including a #102 perf-regression guard at
    100 layers × 100 nodes < 50 ms).
  `pnpm --filter @understand-anything/core test` — 654 / 654 pass.
  `pnpm --filter @understand-anything/dashboard exec tsc -b` — clean.
  `pnpm --filter @understand-anything/dashboard build` — clean.

Pre-existing on master and not from this branch: `pnpm lint` errors
out with "eslint: command not found" — `eslint` isn't installed by any
package and the root `lint` script is bare `eslint .`. Out of scope here.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-04 11:49:20 +08:00