Changelog #
[0.8.0] - 2026-09-27 #
Added #
- Native GitLab Code Quality output through
lintlang scan --format gitlab, with repository-relative locations, stable fingerprints, and a CI example for merge requests and the default branch. Incomplete reports fail explicitly instead of hiding findings behind a successful export. - Parser-backed source locations for structured prompts, tools, schemas, and messages, plus source spans for whole-prompt checks. The shared provenance also improves existing terminal, JSON, and SARIF reports.
[0.7.1] - 2026-09-23 #
Fixed #
lintlang init --githuband the maintained Code Scanning examples now pin the released v0.7.0 Action commit. Existing generated workflows remain untouched unless their owners rerun the initializer with--force.- Dev Container Feature 1.0.1 defaults to the published LintLang 0.7.0 package, including the installer fallback. Current setup guides point to these corrected release routes.
[0.7.0] - 2026-09-23 #
Added #
A conservative
lintlang scan --fixpath for instruction documents applies a bounded in-place rewrite and declines files without a leading instruction section or edits that would introduce an example. Use--dry-runto preview without writing or--backupto preserve the original bytes before writing; automatic repair is not a safety verdict.HERM confidence output now names its coverage drivers, so a low confidence label can be traced to missing or sparse evidence rather than read as a detector accuracy measurement.
Native distribution entry points now include an on-demand Pi scan skill, a GitHub Copilot CLI audit skill, and Cursor marketplace packaging for the portable Claude Code audit skill. Each invokes the standalone scanner on selected files; none scans a whole host session or proves runtime behavior.
A Dev Container Feature installs an exact LintLang release in an isolated environment. Its default remains the published
0.6.0package until the 0.7.0 package is available; the version option can select a later release.A second Agent Plugins 1.0 manifest at the Claude Code plugin root makes its existing skill discoverable by hosts using that format. The automatic Claude Code hook remains a separate surface.
Tool definitions are found by shape, wherever they sit. Previously only a root
tools/functionslist ofname+parameters/input_schemaobjects was read: a root JSON array was an input error, and MCPinputSchematools,mcpServers.<server>.tools, vendor keys such as VS Codecontributes.languageModelTools, name-keyed tool maps and GeminifunctionDeclarationsscannedPASSwith nothing inspected. See "Supported formats and extraction" inllms-full.txtfor the exact signatures and the hard negatives (SBOMs, JSON Schema, OpenAPI, lockfiles, pipeline templates).Every result reports what it inspected (
Inspected:line;inspectedin JSON). A file with nothing to inspect isSKIPPEDwith its reason, neverPASS. A scan in which every file was skipped, and a named file whose tool-like objects could not be read, exit 1;--allow-uninspectedopts out. JSON also gainsnot_inspected,skipped, and a per-findingline.Markdown files are instruction documents, not chat system prompts. YAML front matter with
name/descriptionis read as skill metadata: H1.1/H1.2 on the description, H1.7 (over 1024 characters), H1.8 (no "when to use"), H1.9 (name invalid or different from theSKILL.mddirectory). Front matter is no longer linted as body prose.H4.5: an instruction document references a project file that does not exist.
Findings in text files carry their line number (terminal
file:line, JSONline, SARIF region) and quote the whole offending line.--show-all; the terminal otherwise shows five findings per code and counts the rest. Clean and skipped files in a multi-file scan are summary rows.
Changed #
Terminal scans on a TTY end with one dim package-version and repository pointer. Redirected output, JSON, Markdown, and SARIF omit the pointer.
lintlang init --githubnow generates a workflow pinned to the v0.6.0 Action commit. Existing generated files are reported as different and are not overwritten without--force.Concrete short tool descriptions no longer fail solely on length. One-sided tool-description containment respects distinct input property schemas.
Parameter descriptions are not required to repeat explicit scalar schema or tool-description context. Unexplained parameters still produce H3 findings.
Server manifest instructions no longer inherit host-agent shape requirements. Unresolved localization keys are disclosed as uninspected text, and skill source catalogs are skipped with an explicit reason.
The chat-prompt shape heuristics no longer run on Markdown instruction documents: H5 instruction count without priority ordering, H5 negative density, H6 missing output format, H6 missing version marker, H4 missing boundary vocabulary. Those generic prompt-shape notices did not identify a specific sentence to repair in an instruction document. They still run on
.txt/.promptfiles and config system prompts.H1.3 no longer treats get/set/run/execute/use/make as vague verbs. One-sided H1.6 needs a shared domain term and the same leading verb. H1.4/H1.5/H1.6 compare within one tool container. H3 ignores unions of scalar types.
Python: a string is extracted as a prompt when the code uses it as one; docstrings, help text, log and exception messages are not prompts. P2 is LOW (over 500 characters) / INFO instead of MEDIUM / LOW, and the version-marker note does not apply to extracted literals. A Python tool's schema is counted as inspected only when it is a literal object; dynamic schema expressions are retained as coverage notices rather than silently treated as empty schemas. Python test files excluded during directory walks now appear as explicit SKIPPED results instead of disappearing from the reported coverage.
Baselines: evidence is part of a finding's fingerprint and H2/H4/H5 evidence in text files is now the whole line, so such entries recorded by 0.6.0 resurface once and need re-recording.
Added (earlier in this cycle) #
- On-demand
lintlang-auditskill in the Claude Code plugin (integrations/claude-code/skills/lintlang-audit/), plugin version0.2.0. It audits the file a user names, resolving the released CLI fromPATHor throughuvx --from lintlang==0.6.0, so it needs no checkout of this repository. Previously the plugin shipped only the automaticPostToolUsehook, which covers just the file Claude Code has already changed; there was no way to ask for an audit. The two surfaces are separate: the skill is not a hook, and the plugin README and manifest descriptions now say so. lintlang scan --discover [ROOT]opts in to repository discovery of recognized agent-instruction files anywhere underROOT(default.):AGENTS.md,CLAUDE.md,GEMINI.md,SKILL.md,agent.yaml/.yml/.json,.github/copilot-instructions.md, and*.instructions.mdunder.github/instructions/, matched case-sensitively at any depth — that directory's documented spelling, so a README or a scratch note kept beside the real instruction files is not scanned as one. Editor and host layouts (.cursor/rules,.claude/agents,.windsurfrules) are known omissions, not discovery targets; pass such a file as an explicit argument. Symlinks are not followed, as in a directory scan, but a recognized instruction file skipped for that reason is now named on stderr instead of silently dropping out of the scanned set; an excluded one is not reported. Explicit file arguments remain canonical;--discoverunions its discovered set with them, deduplicated.--excludeglobs and a repository's.lintlangignorefilter discovered files exactly as they filter a directory scan, so a repository that keeps deliberately broken instruction fixtures can keep them out of its own repository-mode gate. Generic directory scanning (lintlang scan <dir>) is unchanged and keeps its own broader, extension-based sweep.lintlang scan - --stdin-filename <virtual-path>scans exactly one document from standard input. The virtual path drives parsing (including.pyAST extraction), reported locations, JSON/SARIF identity, and baseline matching; the path itself is never opened.-without--stdin-filename, more than one-,--stdin-filenamewithout a-input, and-combined with--discoverare all rejected with a usage error (exit 2). One invocation takes one unambiguous source of files: a generator scans what it generated, a repository gate scans the repository. Released 0.6.0 rejected a bare-as a missing file (exit 1), so a wrapper that branches on that exit code sees 2 instead.lintlang init --githubalso recognizes a rootSKILL.mdas an automatic candidate, checked last in the existing detection order so no previously auto-selected repository changes which file it picks. An explicit--pathstill wins over every candidate.
Changed #
- A scan that inspects zero files is an input/coverage error on every
channel. Released 0.6.0 already exited 1 for it and said so on stderr
(
Error: No files were successfully scanned.) and in SARIF, which reportedexecutionSuccessful: falsewith anLL_INPUT_ERRORnotification but noERRORresult; JSON alone still printed[], so the JSON report contradicted the process status. The scan now exits 1 with a matchingERRORresult on every output channel (terminal, JSON, SARIF). New--allow-emptyis the opt-out for a caller that intentionally scans an input that may sometimes be empty: it exits 0 with a stderr note,[]for JSON, and an empty SARIF run reportingexecutionSuccessful: true. Report and exit status now agree in both directions. An interim change onmainafter the 0.6.0 tag made a zero-file scan exit 0; it was never released, so no published version behaved that way and an upgrade from 0.6.0 sees only the added channels and the new flag. The first-party GitHub Action runs the same CLI over itspathinput, so a path that inspects zero files now fails the step with exit 1; the Action exposes no input that forwards--allow-empty, so a workflow that needs the opt-out has to call the CLI directly.--write-baseline's pre-existing empty-scan error, which already refused to write a baseline, is unchanged. - Breaking: the pre-commit hook now consumes pre-commit's own changed-file
selection instead of a hard-coded path.
args: [AGENTS.md],pass_filenames: false, andalways_run: trueare gone from.pre-commit-hooks.yaml; the hook now declares a conservativefiles:regex matching exactly the recognized instruction set above, so it fires only on changed files that are themselves agent instructions, across however many such files changed in one commit. A repository that relied on the hook always running regardless of which files changed should instead runpre-commit run lintlang --all-files. Reproducing the exact old single-path behavior takes all three settings in its own.pre-commit-config.yaml—args: [AGENTS.md],pass_filenames: false, andalways_run: true— because pre-commit appends the changed filenames afterargs, soargs:alone adds a fixed path rather than replacing the selection. The pre-commit guide now states that consequence and the advisory default it sits next to: without--fail-on, a FAIL verdict does not block the commit, because the scan exits 0 whatever it found and only an input error is nonzero. - Behavior change:
--excludeand.lintlangignoreglobs are now anchored and translated correctly. The previous translator rewrote the pattern by sequential string replacement and matched it unanchored, with two consequences:**/became a mandatory rather than an optional path segment, so**/*.mdmatcheddocs/a.mdbut not a root-levela.md; and any pattern matched anywhere in a path, sodocs/**also excludednotdocs/a.mdand*.mdalso excludedmyfoo.md.bak. Both are fixed. Following gitignore, a pattern containing no/still matches at any depth, soCHANGELOG.mdand*.mdkeep excluding nested files. A repository whose exclusion happened to depend on the over-broad matching will now scan those previously skipped files. The same translator backs--discoverfiltering. - H5's per-negative-instruction LOW notices
(
Negative instruction '…' could be reframed positively.) are removed. The aggregated MEDIUM density finding for a prompt with many instructions and no explicit priority ordering is unchanged.
Fixed #
Direct-format tool entries with a non-string
namenow produce a located input error instead of an internal attribute error. The result remainsERRORwith exit code 1.H2 no longer treats a finite
loop overorloop throughinstruction as indefinite without an explicit continuation term. H4 no longer treats a domain invariant such asAlways maintain backward compatibilityas cross-context memory. Their actual unbounded and memory-carrying forms remain findings; a file whose only HIGH or CRITICAL findings were these false positives can now move from FAIL to REVIEW or PASS.Verdict change, in both directions. H2 no longer reports an unbounded-behavior phrase that the prompt forbids.
Do not continue indefinitely; stop at the first terminal result.was reported asUnbounded continuation(CRITICAL) and scanned FAIL; it now carries no H2 finding. One guard covers every H2 unbounded-behavior phrase (keep trying until,retry until,loop until,loop over/loop through,continue until/continue indefinitely). The negator (never,do not,don't,should not,must not, their contractions, andcannot/can'tafter a subject) must sit directly on the phrase, with at most two adverbs from a closed list between them. A file whose only CRITICAL findings were such prohibitions now scans REVIEW or PASS. The guard replaces the narrower negation check thatretry untilhas had since 0.6.0 and is stricter than it: anever retry until …that carries a trailing condition (unless,if,when,except), is asked as a question, is doubly negated, or is split from the phrase by a tab or a blank line is now reported where 0.6.0 was silent, so a file relying on one of those forms can move to FAIL. An interrupted or delegated prohibition (Do not, under any circumstances, …,Do not let the agent …) is still reported; put the negator directly on the phrase. A stop condition in a clause coordinated after the prohibition —Do not continue indefinitely and stop when the queue drains.— does not make the prohibition conditional, so the corrected wording a user writes after being flagged is not flagged again; a condition attached to the forbidden behaviour itself (Do not keep trying until it works when the credentials are wrong.) still is. The guard reads the negator's own clause: to its left a comma opens that clause only when it starts a coordinated one (…, and do not retry until success), so a parenthetical or a complement cannot hide an earlier negative (It is not true, however, that you must never retry until it works.is reported); an earlier negative counts with or without its apostrophe (wontas well aswon't); and a condition introduced after the phrase (…, if the queue is non-empty) defeats the prohibition exactly as an unpunctuated one does, whether the condition is introduced by a comma, a dash, or a parenthesis (Do not retry until it works (if the queue is non-empty).is reported, as its comma form is). A parenthesis that opens anything other than anif/when/whenevercondition is an aside or the author's own bound and leaves the prohibition whole (Do not retry indefinitely (see the runbook).,Do not continue indefinitely (stop after ten items).). Fronted and trailing conditions are treated consistently: bothIf the queue is non-empty, do not retry until it works.andDo not retry until it works, if the queue is non-empty.remain reportable because neither states an unconditional bound. Trailing conditions remain visible through ordinary modifiers (Do not retry until it works, only if the queue is non-empty.), so the prohibition is not mistaken for a bound. Finding descriptions and evidence text are unchanged, so existing baseline entries still match.H4's
Long system prompt with no context boundary markers(MEDIUM) now also requires the prompt to demonstrate cross-context statefulness — carrying state, memory, or history across turns, tasks, or sessions — before it reports. A long, single-shot reference prompt that never asks the agent to carry anything between turns no longer reports this finding on length alone; the other H4 erosion patterns are unchanged. The statefulness signal is a recognizer, so a prompt that asks for carry-over in wording it does not list (Keep the running tally from earlier questions in mind.) is a known miss.H6's
System prompt references multiple output formats (…)(MEDIUM) now requires each named format to carry its own output-format instruction, not a bare mention. A prompt that names two formats only while describing which file types a tool reads no longer reports this finding. An imperative taking the format as a direct object (Always output JSON.,Return JSON only.,Emit XML.) counts as an output instruction; a descriptive third-person clause (The upstream service returns JSON.) does not. That negative is correct only when the clause describes what some other system emits. A two-format delivery of the agent's own reply stated in the third person (The agent's reply is delivered as JSON to the API and as Markdown to the UI.) is recognized as the competing contract this rule exists to surface.H6's
System prompt has no explicit output format specification or example.(LOW) no longer reports on a prompt that does specify one in an ordinary spelling.Return Markdown only.andReturn a plan as Markdown.were both missed because the verb had to be immediately followed byin/as/with/using. The recognizer is deliberately not widened to arbitrary words before the format name, soOutput exactly one Markdown document.is a known remaining miss — as is any wording that puts an unlisted word between the connector and the format name (Reply using the YAML shape below.) or uses an unlisted verb (Produce a single JSON object.). Each documented miss for the narrowed H4 and H6 rules now has a test that asserts the behaviour the rule should have and is expected to fail until the miss is fixed, so closing one of these gaps is a visible decision rather than a silent change.HERM recognizes explicit prose priority statements such as
Priority order is: … then …without treating absence or uncertainty language as an ordering. A public-safe machine-readable case and scanner fixture preserve this boundary.
[0.6.0] - 2026-09-12 #
Added #
- Opt-in baselines for incremental adoption:
scan --write-baseline FILErecords reviewed structural findings;scan --baseline FILEreports and gates remaining findings. Exact repository-relative identities and occurrence counts prevent file-wide or rule-wide suppression. Invalid inputs remain fatal, and baseline creation refuses to overwrite existing paths. - Optional first-party GitHub Action
baselineinput for both terminal and SARIF output, with protection against a SARIF report overwriting the baseline. Reports disclose the acknowledged finding count; HERM scores are unchanged. - The Claude Code hook now resolves LintLang without putting the edited
project's directory on the import path. It prefers the installed
lintlangexecutable, and usespython3 -m lintlangonly with-PandPYTHONSAFEPATH=1on interpreters that support them. Previously the-mprobe ran before the pinned version was compared, so alintlang.pyin an opened project could be executed by the hook. - Root Claude Code marketplace manifest (
.claude-plugin/marketplace.json) cataloging the existingintegrations/claude-codeplugin. Claude Code cannot install a plugin that no marketplace lists, so the adapter previously required a local checkout and--plugin-dir;/plugin marketplace add hermes-labs-ai/lintlangthen/plugin install lintlang@lintlangnow works from the repository. The marketplace is namedlintlangrather thanhermes-labs, which is already taken by the publishedhermes-labs-ai/agent-signagecatalog. - External MegaLinter plugin exposing the pinned LintLang scanner as
AI_LINTLANG, with an exact configuration guide and clean/failing fixture coverage. Verified end to end inoxsecurity/megalinter-python:v9.4.0: the descriptor loads,lintlang==0.5.3installs at run time, a failing fixture exits 1, and a clean-only workspace exits 0.
Fixed #
- H2 recognizes explicit local verification bounds and negated retry prohibitions, reducing false positives for those instruction patterns.
[0.5.3] - 2026-09-02 #
Added #
- Native Gemini CLI
AfterToolextension that returns bounded LintLang repair context after successfulwrite_fileandreplaceedits. - Native Claude Code
PostToolUseplugin for successfulWriteandEditoperations, with shell-safe paths and an exact scanner-version contract. - Native OpenCode 1.18.27
tool.execute.afteradapter that scans explicit changed-file paths and skips ambiguous patch operations.
The Gemini extension passed a copied-install live host run and bundles the repository source with a pinned runtime dependency. All three adapters are non-blocking and never rewrite the changed file.
[0.5.2] - 2026-09-02 #
Added #
- Native Hermes Agent
pre_verifyintegration that scans changed instruction-bearing files once before a coding turn finishes.PASSandREVIEWremain non-blocking;FAILor inputERRORreopens the turn with the exact local scan command.
[0.5.1] - 2026-09-02 #
Fixed #
- Generated GitHub Code Scanning workflows now isolate SARIF upload permission in a second job, disable persisted checkout credentials, and preserve failed scan results through an artifact handoff.
[0.5.0] - 2026-08-24 #
Added #
lintlang init --githubcreates an idempotent, pinned GitHub Code Scanning workflow for an existing repository-owned instruction file. It auto-detects common agent-instruction paths, supports an explicit repository-relative--path, uploads native SARIF, and refuses silent overwrite unless--forceis supplied.
[0.4.1] - 2026-08-15 #
Changed #
- PyPI Homepage and Documentation links now point to the dedicated LintLang product page at https://hermes-labs.ai/lintlang. This is a metadata-only release; LintLang's scan behavior and GitHub Action interface are unchanged.
[0.4.0] - 2026-08-13 #
Added #
- Deterministic SARIF 2.1.0 output for
lintlang scan --format sarif, validated offline against the hash-frozen OASIS errata-01 schema. Findings use their most specific stable code asruleId; severity maps to SARIFlevelwithout security metadata or custom fingerprints. - Repository-relative, URI-encoded artifact locations. AST-extracted Python findings include evidence-supported line spans; structured YAML, JSON, and text findings remain file-level rather than inventing line or column data.
- An optional
sarif-fileinput for the first-party composite Action and a least-privilege example that uploads the generated file with GitHub's immutableupload-sarifAction pin, even when verdict gating returns nonzero.
Changed #
- Release publication now checks out the GitHub Release tag, verifies that its
vX.Y.Zvalue matches the package version and checked-out commit, runs Twine metadata checks, and uses immutable Action SHAs before trusted publishing.
Fixed #
- Quoted detector examples, inline code, fenced code, and metalinguistic descriptions no longer trigger H2, H4, or H5 merely by mentioning detector phrases. Live directives remain reportable, and scope-classification failures preserve prior reporting. This post-0.3.8 fix was merged separately in PR #41 and remains regression-covered.
[0.3.8] - 2026-08-05 #
Added #
H1.6, a sub-code of H1 that reports tool pairs carrying no differentia — where every meaning-bearing term in one description also appears, or has a synonym, in the other. A description states what a tool is; a diagnosis states what distinguishes it from its nearest neighbour. An individually accurate description can still fail to distinguish its tool from a neighbouring tool. H1.6 reports that relational ambiguity before runtime. Schema validation cannot reach it because each colliding tool is individually valid.- Detection uses a curated synonym lexicon, so it reaches some pairs that word
overlap cannot.
Search the documentationandSearch through the docsscore 0.25 on H1.5's Jaccard measure while carrying no distinguishing term under H1.6's model. - Two shapes are distinguished. Mutual — neither tool distinguishes itself. Domination — one tool's every analysed term is covered by the other, so the finding identifies the less-specific description to repair.
Finding.sub_idandFinding.code, so a finding can be cited precisely ("that's an H1.6") without renaming the pattern IDs already in use. JSON output gains acodefield.pattern_idis unchanged.
Changed #
- Some pairs reported in 0.3.2 are now silent by design: a pair whose descriptions reference each other by name is treated as self-disambiguating and skipped, unless one of them declares itself an alias. Diffing findings across versions will show this.
- H1.6 findings are MEDIUM, not HIGH, so they inform a build rather than break
one.
--fail-on failkeys on CRITICAL/HIGH and is unaffected; use--fail-on reviewto gate on them. This is deliberate while the check's recall is unmeasured against a labelled corpus. - H1.6 comparisons are scoped to tool definitions extracted from one parsed input. Directory scans do not aggregate definitions across files or infer that separate files share a selection namespace.
Fixed #
- Descriptions in non-Latin scripts are no longer reported as redundant. Tokenization matched ASCII only, so Chinese, Japanese, Korean, Cyrillic and Arabic descriptions produced no terms at all — and set containment holds vacuously for an empty set, so such a tool read as "dominated by" whatever it sat beside, with advice to delete it. Tokenization is now Unicode-aware, and a tool carrying too little analysable text is skipped rather than compared. The lexicon remains English, so synonyms in other languages are not detected.
- A tool whose description declares it an alias of another is now reported as a
collision. Detection matches a fixed list of phrasings (
Compatibility alias for X,Deprecated. Use X,Superseded by X); the same relationship phrased differently is still missed, the same way the synonym lexicon is finite. - A tool named with an ordinary English word —
access,configure— no longer suppresses findings against its neighbours merely because that word appears in their descriptions. Only an identifier-shaped name counts as one tool naming another.
Note for anyone diffing the source: several entries that appeared here during development described defects introduced and fixed within this unreleased branch, not behaviour any 0.3.2 user encountered. They have been removed. The cardinality and name-suppression problems never shipped.
[0.3.2] - 2026-08-04 #
Added #
- A first-party composite GitHub Action that installs LintLang from the selected action ref and preserves the CLI's verdict-based exit status.
- CI smoke coverage for successful and failing action invocations.
- A native pre-commit hook that visibly reviews explicit repository-owned instruction paths without blocking on heuristic verdicts by default.
Changed #
- The GitHub Actions quick start now uses
hermes-labs-ai/lintlang@v0.3.2. - The quick start now includes exercised
uvxand isolatedpipxpaths. - Repositories can opt into blocking pre-commit
FAILfindings after reviewing their baseline.
[0.3.1] - 2026-07-19 #
Changed #
- Python source scanning uses AST extraction for embedded prompts and the P1/P2 pipeline checks while preserving the offline, deterministic scan contract.
- An unsupported embedding experiment was removed from the candidate before release because an unavailable backend could not be distinguished from a clean result.
Added #
- Provider-neutral preflight candidate: deterministic analysis of one present
instruction plus typed explicit context, with PF001-PF005 exact evidence,
ALLOW | NOTICE | HOLD | UNAVAILABLE | ERROR, redacted-by-default JSON, and explicit source-bound correction previews. It never retrieves history or sends to a provider. - Version-of-record consistency gate (
tests/test_version_consistency.py): assertslintlang.__version__equalspyproject.toml's[project].version, reading source directly so it holds in a fresh clone. Fixes and guards against the prior drift where__version__reported0.2.1while the published artifact was0.2.2. This is the "separate gate" thattest_docs_consistency.pynames as out of its scope. - Fatal input-integrity channel: missing, unreadable, or malformed requested inputs now produce explicit
ERRORresults in the CLI,scan_file(), andscan_directory(), and the CLI exits 1 regardless of--fail-on; another valid input can no longer mask omitted coverage. compute_verdict(result)and the public terminal/Markdown formatters preserve thatERRORstate instead of treating an unread input with zero findings as clean; passing a findings list remains compatible after successful input.- JSON output includes
input_errorfor every path and usesverdict: ERRORfor input failures instead of converting parse errors to INFO/PASS lint findings. - README examples now report 0.3.1, and verdict/detector language is scoped to structural findings rather than runtime guarantees.
- Documentation consistency checks now forbid brittle suite-size claims and scope the bundled samples as regression fixtures rather than accuracy evidence.
Notes #
- Zero runtime dependency change —
pyyamlremains the only runtime dependency. 0.3.0was never published to PyPI or created as a GitHub Release. A publicv0.3.0Git tag already points to an older, pre-fix commit, so this release advances to0.3.1rather than moving or reusing that immutable tag.
[0.2.2] - 2026-04-26 #
Added #
INTENT.mdat repo root — Hermes Labs convention; one-page invariants doc covering accepts/refuses/non-goals + verification contract.evals/sample-detection-rate.sh— runnable regression check that scans the bundled samples and asserts the expected outcomes. This fixture is not an external accuracy evaluation.tests/test_docs_consistency.py— mechanical CI gate (three assertions) that fails the build if the README opener / latest CHANGELOG entry /pytest --collect-onlycount drift apart. Catches the fabrication-class pattern where a chisel pass updates one surface but leaves a stale figure on another. Replaces manual eyeball-grep audits withpip install lintlang && pytest tests/test_docs_consistency.py-checkable invariant.
Changed #
- README refreshed for Hermes Labs Flagship Standard v1. Added detector scope, a comparison with model-based review, explicit non-goals, and a reproduce-yourself line pointing at
evals/sample-detection-rate.sh.
Notes #
- Chisel pass — README + structural docs only. No detector changes.
- Tier B coverage against
flagship-standard.md: 6/7 (B6 plugin path is the acknowledged miss; queued for v0.3 when a formalProtocol/register()extension surface lands). - Experimental E-series detector work was not shipped. Any future integration requires a public evidence corpus, hard-negative tests, and an explicit opt-in contract before it can affect existing CI results.
[0.2.1] - 2026-04-13 #
Added #
- H5 layered exemption system — three-layer filtering reduces false positives on negatives:
- Layer 1: Structural exemptions (HTML comments, code blocks, generated-file markers)
- Layer 2: Phrase-level exemptions (privacy disclaimers, UI labels, descriptive text, idiomatic expressions)
- Layer 3: Safety-context keyword window (existing behavior, now the fallback)
- Expanded vague qualifier detection — catches figurative verbs (
lean into,err on the side of,double down on,keep it simple), broader ambiguous conditionals (if appropriate,when possible) - H6 code-aware format detection — strips fenced code blocks, inline code, filenames, and CLI flags before counting format keywords (prevents
--jsonflag from triggering mixed-format warnings) - Multi-file summary table — box-drawing table with per-file verdict, findings breakdown, and scan timing (terminal output only, shown when >1 file scanned)
- Vague qualifier deduplication — identical matches within a file are reported once
Changed #
- Development status upgraded from Alpha to Production/Stable
- Author metadata updated.
[0.2.0] - 2026-03-25 #
Changed #
- Breaking: Replaced numeric HERM score with PASS/REVIEW/FAIL verdict in terminal and markdown output
- ❌ FAIL — any CRITICAL or HIGH finding
- ⚠️ REVIEW — any MEDIUM finding
- ✅ PASS — only LOW/INFO findings or none
- Terminal output now leads with verdict + severity summary instead of dimension bars
- Markdown report restructured around verdict + findings (no score in header)
- JSON output: verdict at top level, HERM score moved under
hermkey (preserved for programmatic use) patternscommand simplified to show H1-H7 detectors only
Added #
--fail-on fail|reviewCLI flag for verdict-based CI gatingcompute_verdict()function in public APItest_verdict.pywith 10 dedicated verdict logic tests.mdextension support inscan_directory(SKILL.md files were silently skipped)- Expanded
is_prompt_likeregex to recognize SKILL.md format (description/purpose/role patterns)
Fixed #
- SKILL.md files now get proper coverage instead of defaulting to 65% (low confidence)
- Scanning directories with .md instruction files now includes them automatically
Deprecated #
--fail-under(HERM score threshold) still works but--fail-onis preferred
[0.1.2] - 2026-03-02 #
Changed #
- Updated project URLs for PyPI backlinks (Homepage, Documentation, Repository, Bug Tracker, Changelog)
[0.1.1] - 2026-03-02 #
Fixed #
- Standardized package metadata for Hermes Labs.
- Fixed publish workflow to use API token authentication
- Added community health files (CONTRIBUTING.md, SECURITY.md, CODE_OF_CONDUCT.md)
- Added dependabot configuration
[0.1.0] — 2026-03-01 #
First public release.
Core #
- HERM v1.1 scoring engine (6 dimensions, 8 signal categories, coverage/confidence)
- H1-H7 structural detectors with Finding dataclass
- YAML, JSON, and plain text parsers with auto-detection
- Terminal (ANSI), Markdown, and JSON output formats
--fail-underflag for CI gating
CLI #
lintlang scan— scan files or directorieslintlang patterns— list available patterns and dimensionspython -m lintlangsupport via__main__.py--format,--patterns,--min-severity,--no-suggestionsflags- Dynamic pattern choices from registry
Detectors #
- H1: Empty/short/vague tool descriptions, duplicate names, word overlap (Jaccard + stopwords)
- H2: Missing constraint scaffolding, unbounded retry loops
- H3: Phantom required fields, missing param descriptions, generic names, nested object inspection
- H4: Context boundary erosion, missing scope signals
- H5: Negative instruction density, vague qualifiers
- H6: Mixed output formats, missing format specs, template variable detection
- H7: System message placement, consecutive roles, orphan tool results
Programmatic API #
scan_file(),scan_directory(),scan_config()ScanResult,HermResult,AgentConfig,Finding,Severityexports- PEP 561
py.typedmarker