P1: Uncalibrated Threshold #
A threshold-like numeric assignment lacks a nearby recognized calibration or justification comment.
Applies to: Python numeric assignments recognized as thresholds. Manual repair; no automatic fix for this example.
What triggers it #
A recognized nonzero threshold assignment without a calibration comment emits MEDIUM. A recognized comment with uncertainty markers can emit LOW. The scanner inspects the comment signal, not the validity of the calibration.
How to repair it #
Document the actual basis for the threshold near the assignment. Include the policy or measured evidence that another maintainer can reproduce.
Reproduce the finding #
Use uv and Python 3.10+, plus curl. Run the examples in a scratch directory.
confidence_threshold = 0.8
curl -fsS https://lintlang.ai/examples/rules/p1-bad.py -o p1-bad.py
uvx --from lintlang==0.8.0 lintlang scan p1-bad.py --format json
Expected with 0.8.0: the JSON report includes P1, severity MEDIUM. This is an illustrative policy threshold: four of five equally weighted checks. The arithmetic is reproducible; it is not an empirical model-confidence claim. Use your real policy or evaluation evidence in production.
Improved example #
Download the improved example.
# Derived from this illustrative policy: accept when at least 4 of 5 checks pass.
# Five equally weighted checks give a measured score of 4 / 5 = 0.8.
confidence_threshold = 0.8
curl -fsS https://lintlang.ai/examples/rules/p1-improved.py -o p1-improved.py
uvx --from lintlang==0.8.0 lintlang scan p1-improved.py --format json
Expected with 0.8.0: P1 is absent. Other diagnostics may still appear; this repair targets the rule above.
Both commands use advisory mode: a finding does not itself make the command fail. To fail CI on MEDIUM findings, add --fail-on review. --fail-on fail gates only HIGH and CRITICAL. See outputs and exit codes.
Detection details #
Released P1 detection contract and scope
Python source is parsed, not imported or executed. The extractor recognizes
string literals of at least 50 characters with a supported prompt signal.
F-string literal portions are retained with {...} placeholders for expressions;
runtime interpolation and data flow are not evaluated. Prompt candidates are
deduplicated by their first 200 characters, so distinct strings with the same
prefix can collapse to one candidate. Python syntax errors produce ERROR.
Only H2, H4, H5, and H6 run on extracted prompts. Literal Python tool definitions provide a separate declared surface on which H1 and H3 run; H7 does not run in Python extraction mode. Selecting a family for which a Python input provides no supported surface is not proof that the check passed. P1/P2 still run independently of that selection.
- P1: Uncalibrated Threshold. Inspects supported nonzero numeric assignments to threshold/confidence-like names, excluding recognized counters. No recognized nearby calibration comment produces MEDIUM; a recognized comment containing uncertainty markers can produce LOW. This tests a documented justification signal, not whether the threshold was actually calibrated or is correct.
- A string is extracted as a prompt when the code uses it as one: it is bound to
a prompt-like name, keyword argument or message-dict key (
prompt,system,instructions,template,content, ...), or it addresses a model ("You are", "your task"), or it matches three or more prompt signals. Docstrings, bare string statements, and text passed to logging, argparse/click help, or an exception are never prompts. - Tool definitions written as literals are read: a call or dict literal with a
literal
name, a schema keyword (inputSchema,input_schema,parameters,parameters_json_schema,args_schema) and a literal or absent description — the way Python MCP servers declareTool(name=..., description=..., inputSchema=...). H1 and H3 run on them, with the call's line. Tools declared by decorator and docstring, or with a computed description, are not read. A dynamic or non-object schema expression is not evaluated: the literal name and description remain inspected, while the schema is excluded fromtools_with_schemaand named undernot_inspected. - P2: Embedded Scaffold. An extracted prompt longer than 500 characters produces LOW; one longer than 200 and at most 500 produces INFO. It suggests reviewing whether the prompt should be externalized, not that embedded prompts necessarily cause runtime failures.
A Python file with no recognized prompt may still have P1 findings. A clean scan cannot establish that dynamically assembled prompts or every pipeline were inspected. Python decoding currently ignores invalid UTF-8 bytes; ordinary text/configuration loading uses strict UTF-8. Neither path executes source code.