Grasp ยท Comprehension Index Scan Report

Who still understands this codebase?

A metadata-only scan of cline/cline measuring how much of the code any currently-active human has genuinely written or engaged with โ€” and where comprehension debt is concentrating.

Repository github.com/cline/cline Scanned July 22, 2026 History window 3 years ยท 6,515 non-merge commits Engine v0.2 ยท line-weighted, full-history scan ยท git metadata + sampled API review signals
Headline finding

The evals/ module is dark code โ€” and it's still hot.

53% of its lines are AI-attributed and 45% were written by contributors no longer active. Only 2% of the module was authored by anyone who has committed in the last 90 days โ€” no single active contributor holds even a 10% share โ€” yet it received 129 commits in the last 180 days. In our framework, sustained change without living commit-authorship is a leading risk indicator for incident cost and onboarding drag โ€” an indicator, not a verdict on any contributor.

Living knowledge
79.2%
of recent lines authored by currently-active humans (line-weighted)
Contributor survival
35 / 322
humans active in last 90 days vs. all-time authors
AI-attributed floor
โ‰ฅ 5.0%
lines with hard AI authorship evidence; true share is higher
Departed-author share
15.7%
lines whose authors are no longer active anywhere in the repo

Dark-code ranking

Modules ranked by comprehension risk: how much of the module lacks a living author, amplified when the code is still changing (heat) and thinly held (bus factor). The green meter is the share written by currently-active humans.

ModuleRiskLiving knowledge AI-attr.Departed Bus factorHeat (commits/180d)
evalsCRIT 98
2%
53%45%0129
localesCRIT 89
1%
0%99%030
protoCRIT 86
10%
5%85%042
.changesetHIGH 72
28%
7%64%094
src/standaloneHIGH 72
6%
0%94%05
scriptsHIGH 70
28%
7%66%143
docsMED 58
37%
18%45%2500
.clinerulesMED 56
32%
37%31%229
src/sharedMED 54
41%
3%56%2198
src/apiMED 49
35%
8%57%10
src/coreLOW 48
48%
6%46%2736
src/servicesLOW 48
48%
4%48%2134

Modules under 200 changed lines in the window are omitted. src/core scores LOW despite 736 recent commits because roughly half its lines have living authors and knowledge is spread across multiple active contributors โ€” heat without darkness is healthy activity.

Review depth โ€” who actually looked at this code?

API-sampled signals joining PR timing with diff sizes from git history: 439 merged PRs timed, with review detail on a size-stratified sample of 34 deliberately weighted toward the largest merges โ€” these figures describe that sample, not all PRs, and describe process patterns, not the diligence of any individual contributor.

No recorded human review
41%
of sampled merged PRs show no human review event in the API record
AI-reviewed only
35%
show reviews from AI reviewer bots (ellipsis, greptile) and none from a human
Comment-free approvals
32%
of human-reviewed sampled PRs were approved with no written review comment
Large fast merges
5
sampled PRs of 300+ lines merged in under 30 minutes

Largest fast merge in the sample: PR #11862 โ€” 2,077 lines touching apps/ and sdk/, merged 24 minutes after opening. Baseline latencies scale sanely with size (median 34 min for โ‰ค50-line PRs, 31 hours for 1,000+ lines), which makes the exceptions the signal, not the norm. Note that legitimate workflows produce these patterns too โ€” pre-coordinated changes, pairing, generated files, release automation, and review conducted outside GitHub's review feature are all invisible to this record.

Contrast scan โ€” two repos, two failure modes

The same engine run on pallets/flask (16 years old, pre-AI era, 3,812 commits) shows the Index distinguishes different kinds of knowledge risk rather than flattering old code or damning new code.

cline/cline (AI-era, 2 yrs)pallets/flask (pre-AI, 16 yrs)
Living knowledge (weighted)79.2%83.8%
AI-attributed floor5.0%0.1%
Active humans (90d)35 of 322 ever1 of 866 ever
Worst moduleevals โ€” 2% living, hotsrc/flask โ€” 51% living, quiet
Dominant risk shapeComprehension debt: fast-growing code with no living author, where sampled merges often show comment-free or AI-only recorded reviewConcentration risk: knowledge is alive but every module is bus-factor 1 โ€” a single active contributor holds the dominant commit share

Flask's headline numbers look healthier โ€” until you see that its living knowledge is concentrated in a single active contributor. The two repos carry risk in opposite shapes, and a velocity dashboard would flag neither.

State of Comprehension Debt โ€” 22-repo benchmark

The same engine run across 22 public repositories in five cohorts: AI-era agents and apps, AI vendor SDKs, modern human-led projects, and pre-AI controls. Sorted darkest first by living knowledge.

Lowest scores
3โ€“5%
living commit-authorship in the OpenAI & Anthropic Python SDK repos โ€” both built by automated codegen pipelines (see note)
Corpus median
61%
living knowledge across all 22 repos
Healthiest repo
94%
curl โ€” 25 years old, 55 active humans. Age isn't the risk; maintenance is.
This repo
79%
cline/cline ranks 8th of 22 โ€” healthy overall, with concentrated dark corners
RepositoryCohortLiving knowledge AI-attr. floorDepartedActive / all-time humans
openai/openai-pythonAI SDK
3%
64.5%32%7 / 158
anthropics/anthropic-sdk-pythonAI SDK
5%
64.0%31%11 / 62
gin-gonic/ginpre-AI control
11%
19.9%69%13 / 540
expressjs/expresspre-AI control
15%
0.0%85%10 / 389
All-Hands-AI/OpenHandsAI-era agent
22%
48.8%29%38 / 499
langchain-ai/langchainAI-era app
24%
1.1%75%41 / 3,683
langgenius/difyAI-era app
32%
36.4%32%167 / 1,400
crewAIInc/crewAIAI-era agent
48%
39.7%12%28 / 310
sst/opencodeAI-era agent
53%
30.1%17%147 / 989
continuedev/continueAI-era agent
55%
3.6%41%4 / 534
browser-use/browser-useAI-era agent
60%
3.5%37%17 / 355
django/djangopre-AI control
62%
1.6%37%59 / 3,409
RooCodeInc/Roo-CodeAI-era agent
62%
15.5%23%22 / 308
fastapi/fastapimodern human-led
69%
12.0%19%12 / 907
vercel/aiAI SDK
72%
15.6%13%88 / 681
lobehub/lobe-chatAI-era app
73%
23.8%3%42 / 354
tailwindlabs/tailwindcssmodern human-led
78%
1.1%21%23 / 372
cline/cline ยท this reportAI-era agent
79%
5.0%16%35 / 322
astral-sh/ruffmodern human-led
83%
0.4%16%111 / 945
pallets/flaskpre-AI control
84%
0.1%16%1 / 866
psf/requestspre-AI control
86%
0.1%14%12 / 792
curl/curlpre-AI control
94%
0.0%6%55 / 1,579

What the corpus shows. Cohort medians: modern human-led 78% ยท pre-AI controls 73% ยท AI-era agents 55% ยท AI-era apps 32% ยท AI SDK cohort 5% (two of its three members โ€” the OpenAI and Anthropic Python SDKs โ€” score 3% and 5%; the third, vercel/ai, is largely hand-written and scores 72%). An essential reading note on the SDK repos: both are produced by automated code-generation pipelines from API specifications โ€” a deliberate, industry-standard engineering choice. Low living commit-authorship in a generated repo reflects that workflow; the relevant human comprehension plausibly resides in the generator and the API specification, which repository-level metrics structurally cannot observe. These scores describe public commit history only โ€” they are not a claim that any organization's engineers do not understand their products, and not a claim about the quality, security, or fitness of any software. Elsewhere the pattern is starker: comprehension debt predates AI. Express โ€” 15 years old and foundational to the npm ecosystem โ€” shows 85% of its recent lines authored by contributors no longer active, while 25-year-old curl is the healthiest repository in the corpus. Age determines nothing; living maintenance does.

Two weighting modes, one conclusion

The engine measures living knowledge two ways, and the two proxies err in opposite directions โ€” which is why we run both. Line-weighting (this report) weights authorship by lines added: it matches "what fraction of the code by volume has a living author," but lets verbose authorship โ€” especially generated code โ€” dominate. Touch-weighting weights each file-touching commit equally: it better matches how comprehension actually forms (repeated engagement, not text volume), but counts a 3,000-line generated commit the same as a one-line fix. Scores are comparable only within a mode, and every Grasp report states which mode produced it. The corpus was computed in both modes, and the findings hold in both: the codegen SDK repos score 3โ€“5% line-weighted and 1โ€“4% touch-weighted; curl scores 94% and 95%; the cohort gradient is unchanged. When two differently-biased instruments agree, the conclusion is stronger than either alone.

How this is computed

  • Living knowledge โ€” share of lines (3-year window) authored by humans who committed within the last 90 days.
  • AI-attributed โ€” lines from commits authored or co-authored by identifiable AI agents (Copilot, Claude, Cline's evaluation agent, bot accounts). Hard evidence only.
  • Bus factor โ€” number of active humans holding โ‰ฅ10% of the module's human-written lines.
  • Risk โ€” the unknown share (1 โˆ’ living knowledge), amplified by recent change rate and thin ownership.
  • Automation bots (CI, dependabot, release tooling) are excluded entirely. Lockfiles, build output, and binary assets are skipped.

Honest limitations

  • This is an open-source repo, so "departed" includes drive-by contributors โ€” it overstates attrition relative to a company repo, where the roster would come from HR data.
  • AI attribution is a floor: squash merges strip co-author trailers, so true AI share is materially higher than 5%.
  • Review-depth figures come from a 34-PR stratified sample (unauthenticated API budget); an authenticated scan would cover every PR.
  • The Index measures engagement, a proxy for comprehension โ€” treat it as a potential indicator, not a diagnosis.

Scope of claims

All data in this report derives from publicly available repository metadata (commit history and pull-request records) analyzed with the disclosed method. Every metric is a defined proxy computed from that record; living commit-authorship is an indicator of comprehension risk, not a measurement of any person's or organization's understanding. Nothing here asserts misconduct, negligence, or lack of competence by any contributor, maintainer, or company, and nothing here is a claim about the quality, security, or fitness for purpose of any software. Repository and project names identify public data sources; all trademarks belong to their owners, and no affiliation or endorsement is implied. Believe a number is wrong? Contact support@graspscore.com โ€” we correct verified errors and version this report.