Python Code-Change Analyzer
URL slug: python-code-change
Takes a unified git diff covering Python (.py) files and turns each changed
function, async function, method or class into its own analysis unit — enough
for an LLM to catch a route added without an auth dependency, a query that
forgot its tenant scope, or a blocking call dropped into an async path.
Parsing strategy
Section titled “Parsing strategy”PythonCodeChangeParser (opentremor_analyzer_python_code_change/parser.py)
streams the diff hunk by hunk:
- Walk the diff, yielding one hunk at a time.
diff --git a/... b/...headers mark file boundaries;.pyfiles are kept, anything else is dropped before a single hunk line is looked at. - Within a hunk, find block-opening lines (
def,async def,class) and locate where each one ends by indentation — the first non-blank line indented at or below the opener. Blank lines never end a block. - Absorb decorators. A run of
@decoratorlines immediately above an opener becomes part of that unit; dropping them would discard exactly the evidence the auth rule reads. - Keep only blocks with a real change — a block containing nothing but context (
-prefixed) lines is discarded. Every+/-line no block claimed becomes onemoduleunit per hunk.
Why not ast?
Section titled “Why not ast?”A diff hunk is a fragment. It is routinely unbalanced, starts mid-function, and
is missing the class statement that encloses it, so ast.parse raises
SyntaxError on most real input. Regex and indentation over the diff text is
what actually survives a hunk.
Block-to-identity mapping
Section titled “Block-to-identity mapping”| Block opens with | type | name |
|---|---|---|
def load_settings(path): at column 0 | function | load_settings |
async def fetch(url): at column 0 | async_function | fetch |
def reset_password(self, ...) indented | method | MemberService.reset_password |
async def run(self, ...) indented | async_method | JobRunner.run |
class Repository: | class | Repository |
| changed imports / module constants | module | (module-level) |
A class’s methods are indented deeper than the class statement, so they are
absorbed into the class’s unit rather than emitted alongside it — one unit
per top-level block, whatever that block contains.
Where the class name in a method’s name comes from
Section titled “Where the class name in a method’s name comes from”When a hunk starts mid-class, the class statement is above the hunk and its
name is nowhere in the diff body. Git puts it in the hunk’s section heading —
the text after the second @@:
@@ -40,7 +40,7 @@ class MemberService: def reset_password(self, user_id):That heading is the only place the enclosing class survives, so it is what qualifies the unit’s name. When git emits no heading, the method’s name is left unqualified rather than guessed at.
Metadata
Section titled “Metadata”metadata.source_path— the changed file’s path, from the diff header.metadata.group_path— the containing Python package (src/opentremor_core/controllers/routers/admin.py→opentremor_core.controllers.routers), filling the same slot a Terramate stack path fills for other analyzers.
Routing
Section titled “Routing”file_globs is ["*.py"], so auto-routing sends any .py file in a diff here
by path glob. sniff() is deliberately not overridden: ingest() only
understands unified-diff input, so a bare .py file pasted into the Analyze
form would sniff as Python and then ingest to zero units. Opting out of
content-sniffing is the honest answer, and diff-shaped input routes correctly
regardless.
The glob is broad on purpose for a first release — it claims test files,
migrations and __init__.py alongside everything else. Narrowing it should be
driven by real false-positive data rather than guessed at, and an organization
that disagrees can disable the analyzer outright with the per-org toggle, with
no code change.
One implicit variant (rules/python-code-change-rules.md); list_rule_types()
is empty. Ten rules across six categories:
| Rule | Severity | What it catches |
|---|---|---|
PYCC-001 | CRITICAL | Tenant-owned data read or written without an org_id scope |
PYCC-002 | CRITICAL | A route added or changed without the auth dependency its neighbours have |
PYCC-003 | HIGH | A secret, key or credentialed connection string in plain text |
PYCC-004 | HIGH | A blocking call added to an async def |
PYCC-005 | HIGH | shell=True, eval/exec, or unsafe deserialization |
PYCC-006 | MEDIUM | A new outbound call with no timeout |
PYCC-007 | MEDIUM | An exception handler that swallows without logging or re-raising |
PYCC-008 | MEDIUM | An import crossing an architectural boundary |
PYCC-009 | MEDIUM | A comment or docstring that now contradicts the code beside it |
PYCC-000 | INFO | Nothing found |
Like every analyzer’s built-in ruleset, these seed into an organization’s
custom_rules as editable rows at org setup — they can be reworded, disabled
or replaced per organization without touching this package.
PYCC-009 is the rule with no static-analysis equivalent at all, and the one
most worth watching: it ships at MEDIUM/LOW confidence so that it never
gates anything while its false-positive rate is still unknown.
Two constraints in the ruleset apply to every finding, not to any one rule.
A finding reports on the code, never on the author — no speculation about
why a change was made or whether it was meant well. An accusation is
unfalsifiable, so a reviewer cannot check it the way they can check “this query
has no org_id”, and being wrong about intent costs far more trust than being
wrong about a line of code. And a title belongs to its rule_id: the titles
are labels for the rules, not a free vocabulary, so a change no rule covers is
PYCC-000, never the closest-sounding label. A finding wearing a title whose
rule it does not match is unreviewable — the reader checks it against the wrong
definition.