Codebook and Method: University AI Task Force Reports

University AI Task Force Reports: Inventory and Codebook

v1.5.1 · R1, R2 and selective liberal arts institutions · 9 September 2026 · Kyle Saunders

What this is

An inventory of deliberative reports on AI produced by named institutional bodies (task forces, committees, working groups, commissions) at Carnegie Research 1 and Research 2 universities and at 46 selective liberal arts colleges. The unit of analysis is the report, not the institution.

Why it exists

Several collections catalogue university AI policies. None catalogues the reports. Verified 2026-09-01:

Collection Size Last updated Why it doesn’t cover this
Eaton, Institutional AI Policies & Governance Structures 19 rows 2024-11-03 No date field; “Policy Link” is a label, not a URL; includes a K-12 district and a national library
Eaton, Syllabi AI Policy Repository 217 entries / 188 institutions 2026-07-17 (alive) No URL column at all; unit is the course
HESA AI Observatory, Policies & Guidelines not stated 404, dead Blog series ended ~May 2025
Mendolia Padlet ~129 posts ~June 2026 No dates, no genre, ownership changed to Western Univ. of Health Sciences
eduaipolicy.org tracker 3,195 sources / 828 institutions 2026-08 sourceType has 4 values, none is “report”; 14/3,195 match task-force patterns; 827/828 agent_reviewed
Teleki et al., FAccT ’26 (ACAI-US79) 79 institutions 2026-06 Its task-force column holds 70 links across 41 institutions, and only 2 are actual reports
arXiv 2406.18842 / Nature Hum Behav 80 guidelines, 28 countries corpus closed 2024-04 “Task force” appears zero times; ~8 of 80 are government documents

Across 24 document-analysis studies checked, a task force report is the sampling unit in none. The two fields no existing collection carries are document date and convening authority. Both are first-class here.

Coverage: all 187 R1, all 139 R2 and 46 selective liberal arts colleges

Headline set: 99 deliberative reports: 84 at 63 R1 campuses, 14 at 12 R2 campuses, and 1 at 1 selective liberal arts college. By access: full text 89, public summary of a gated report 6, abridged 2, one-page abstract 2. By genre: report 84, committee guidelines with recommendations 14, strategy plan 1. Supplementary, excluded from counts: 26 rows. Nine carried over from v1.4.4: UC systemwide Academic Senate; Dartmouth, single-author adviser report; Ohio State Fisher, slide deck; Alabama Graduate School committee guideline, college scope; MIT scholarly-content whitepaper, companion without recommendations; James Madison spring 2024 interim update; UNLV 2024 Faculty Senate report, now behind Google sign-in; Amherst 2023 task force report, behind community login; Colorado College GenAI philosophy, principles without findings. Two moved from the headline set under the inclusion rule: in v1.5, CU Denver’s AI Faculty Liaison Community of Practice, a programme convened by three faculty fellows rather than a chartered body, whose recommendations run to units; in v1.5.1, Boston College’s Student AI Advisors report, a staff-authored account of a student advisory group convened by two offices, whose recommendations run to faculty practice. Eighteen added from the v0.5.0 triage: charge documents for bodies whose report is not yet posted (NC State, New Mexico State, UC Riverside, Appalachian State); principles and values statements (New Mexico, Colorado College); survey and retreat reports (Cal State Long Beach, Delaware); guidance an office or council issued to instructors (UT Arlington, Old Dominion, Kennesaw State, Clark Atlanta); a one-page plan (Marshall); a grant project team’s report (Georgetown); the UC systemwide 2021 working group report as posted by UCSF; and two gated reports (Villanova; UNLV). Three of the eighteen duplicated rows already present (Amherst, Colorado College, UNLV) and were merged in v1.5, leaving fifteen new rows.

Denominator: the 2025 Carnegie Research Activity Designations file, by UnitID. The file lists 187 Research 1 and 139 Research 2 institutions. v1.0 and v1.1 built the R1 list from IPEDS C21BASIC = 15 plus name-matching of the 2025 additions; matching against the Carnegie file itself found two errors, both corrected in v1.2: University of Alabama in Huntsville was R1 under the 2021 classification and is R2 in 2025 (145 of the 146 carried over, not all 146), and Weill Cornell Medical College is R1 in 2025 and was missed. Fifteen institutions in the census are HBCUs: Howard (R1), thirteen R2s, and Spelman in the liberal arts segment. Per-institution status, UnitID, class, HBCU flag, crawl metrics and map join are in crawl-coverage-v1.csv.

status R1, v1.0 (v0.2) R1, v1.2 (v0.4.2) R2, v1.2 (v0.4.2) R1, v1.5 (v0.5.0) R2, v1.5 (v0.5.0) LAC, v1.5 (v0.5.0)
Report found 55 63 10 64 12 1
Crawled; no public AI report surfaced 109 121 120 116 115 43
Site blocked automated access 11 2 1 3 1 1
Crawl too shallow to say 11 1 8 4 11 1
Not crawled (domain mapping) 1 0 0 0 0 0

Status is assigned mechanically: blocked when fewer than 30 HTML pages were read and at least 40% of requests were refused (403, 429, 503); too shallow when fewer than 60 pages were read for any other reason; found when a headline row joins to the institution’s UnitID. Across the 372 census rows at v1.5: 77 found, 274 no report surfaced, 16 too shallow, 5 blocked. The too-shallow line grew from 9 to 16 when the census metrics were replaced with the v0.5.0 values, because the 8 September pass read fewer than 60 pages at seven sites the v0.4.2 pass had read normally (North Dakota, Texas State, Texas Southern and Albert Einstein returned no HTML at all; Arizona State, NYU and Houston stopped under 40). Those seven are protocol failures on the day, not findings about the sites.

Convening authority, headline set: read from the documents in v1.4; the distribution is in the Coded dimensions section below, which supersedes the landing-page values this line used to carry.

HBCUs. None of the fifteen HBCUs in the census has a findable report. Under the v0.5.0 pass, nine R2 HBCUs, Howard and Spelman were read normally (Howard University 142, Clark Atlanta University 109, Delaware State University 146, Hampton University 141, Jackson State University 109, Morgan State University 126, North Carolina A & T State University 292, Southern University and A & M College 199, Spelman College 144, Tennessee State University 185, Virginia State University 141 pages) and surfaced no AI deliberation; four (Florida Agricultural and Mechanical University 20 pages, Prairie View A & M University 8, South Carolina State University 11, Texas Southern University 0) were too thin to say. Three of those four had been read normally under v0.4.2 (103, 103 and 150 pages), so the thinness is the 8 September crawl’s and not the sites’. Clark Atlanta’s AI Task Force guidelines of August 2024, surfaced by v0.5.0, are a supplementary row: syllabus language rather than recommendations to the institution. Given the recall figures below, read this as “no findable report,” not “no deliberation.”

R2 as a population. Fourteen reports at twelve of 139 institutions, against 63 of 187 at R1. Two cautions before reading that as a 9% versus 34% gap. First, R2 recall is unmeasured. Earlier versions said that no R2 report was known before the crawl, and that was wrong: RIT’s April 2024 report was harvested on 1 September 2026 and sat in the supplementary set under an R1-only rule that v1.2 had already retired, four days before the 5 September R2 crawl. It was a known positive that an outdated rule excluded, not a report the harvest lacked, and the v0.4.2 pass read 146 pages at RIT without surfacing it. A single known positive is too few to estimate recall from, and that is the limitation that stands: there is no held-out set for R2, and R2 governance sites are thinner and less uniform (the first R2 pass exhausted its queue at 52 of 139 sites before the crawler learned to follow governance offices reached by path rather than subdomain). Second, the tuned R2 pass produced 31 triaged candidates at 19 institutions, the set published in crawl_out_r2/triage_candidates.json. The first-pass figure carried in earlier versions, 26 candidates at 15 institutions, is not reproducible from the published outputs: no triage artifact survives for that pass, and applying the same score threshold to the v0.4.1 crawl outputs yields 31 candidates at 17 institutions before deduplication. Either way the tuning did not change the picture much. Differential retrieval may affect the observed gap, and neither the size nor the direction of that effect is established here.

Systemwide documents. Indiana’s March 2024 report is a University Faculty Council document covering all IU campuses; it is attributed to Bloomington for the join and Indianapolis carries a note. The UC systemwide Senate report is supplementary because it is not a campus. A many-to-many report–institution mapping remains planned.

Recall, measured. The crawler is run blind against the 55 R1 institutions whose reports were known before the v1.1 re-run (search harvest plus external audit), a holdout frozen in bench_holdout.json. Crawler v0.5.0 re-finds reports at 44 of 55 institutions (80%) and 53 of 69 documents (77%). The v0.4.2 crawler used from v1.1 through v1.4.4 scored 42 of 55 (76%) and 50 of 69 (72%) on the same test, and the v0.2 crawler 26 of 55 (47%; 29 of 69 documents, 42%), under the same matching rule. The two institutions v0.5.0 recovers are Minnesota, whose reports sit in an institutional repository, and UC Davis. bench.py crawl_out_v050/slice*.json and bench.py crawl_out_v042/slice*.json reproduce the three figures. The “roughly 55–60%” stated in v1.0 was an estimate from agent notes and was too generous. bench.py reproduces the figures from the published crawl outputs, against a held-out set frozen in bench_holdout.json and flagged as bench_holdout in the inventory CSV, because the v1.4 coding pass rewrote the column the held-out set used to be derived from; a match is an identical normalized URL, a Drive file ID, or a document basename at 85% similarity, so a re-uploaded file with a -1 suffix still counts. Misses, by cause: automated access refused (UC Davis and Princeton most heavily; Kentucky is the only host that rate-limited, two 429s; Minnesota’s repository refuses datacenter IPs and opens in a browser; Michigan was a miss under v0.4.1 and is found under v0.4.2); the report is no longer linked from any governance page (UNC-Chapel Hill chancellor site, Penn State AI hub, Kentucky ombud office, UTEP AI hub); the linking page is script-rendered (Buffalo); the report lives on a host no governance probe reaches (Penn’s Price Lab, Boston College’s teaching center, Northeastern’s and USC’s senate uploads). A link-following crawler cannot recover documents that are no longer linked; search engines found those because they indexed them earlier. Every negative should still be read as “not linked from the governance surfaces reached, under this protocol.”

Crawler changes, v0.2 → v0.4.2, each made after diagnosing a specific miss or census failure: priority queue (AI-tagged links, then committee and report index pages, then other governance pages) replacing first-in-first-out, because the request budget was being spent on breadth before reaching archives (Michigan Tech); per-section and per-host page caps, because AI hub subdomains absorbed the whole budget (Buffalo, UNC, Penn State, Northeastern); relative links resolved against the URL actually served after redirects (UTEP); off-domain document capture for Drive, Box, SharePoint and repositories (Michigan, UNMC, UNLV, UCSB); corrected seed domains where the IPEDS web address names a sub-site (Minnesota, Hawaii, MIT, Indiana, URI, MUSC, Rutgers) and alternate roots (Washington’s uw.edu, Indiana’s indiana.edu, New Orleans’s lsuneworleans.edu); 308 redirects followed (Drexel); browser user-agent fallback on 403 and curl fallback on TLS handshake failure (UTSA), both logged per host; 429 backoff; a cookie jar (Fordham’s CAS gateway); “www.” prefixed only when it resolves (Oklahoma State’s medical campus); governance offices followed by path from the homepage (52 thin R2 crawls); file-extension skips anchored to the end of the path (“.js” was matching “.jsums.edu” and blanked Jackson State); fifty subdomain probes and sixty path seeds; AI-tagged news pages admitted, at most fifteen; request budget 200, extended to 350 when nothing AI-tagged has surfaced. Depth was left at three: no diagnosed miss was a depth problem.

Crawler v0.5.0 (8 September 2026), written after a reader reported Boise State. The v0.4.2 pass had fetched Boise State’s three AI Coordinating Council white papers and rejected them. Three faults, each fixed. The BODY and OUT regexes ended in \b, so no plural matched: “reports”, “committees”, “task forces”, “councils”, “recommendations”, “white papers” and “guidelines” all scored zero. Google Docs and SharePoint URLs earned no document points because the scorer looked for a file extension. A body named on a page lent nothing to the documents that page linked, so a white paper titled without the council’s name scored as an orphan. v0.5.0 matches plurals, scores hosted-document URLs by host, and lets a page that names a deliberative body pass that context to every document it links; a guard excludes CDN and social hosts from the candidate pile. The precision cost is real: page-level inheritance also promotes bibliography entries, vendor documentation and external frameworks, so the candidate pile is noisier and triage is now the bottleneck. On the full census the count of AI candidate links went from 720 to 1,752, and 66 institutions went from zero candidates to some.

The archive-without-AI pattern, restated for what the metric measures. The rule: no AI report found, and twenty-five or more links reached that score like task force or committee documents. Under the v0.5.0 census, twenty-four R1 and twenty-two R2 institutions meet it (forty-six; three liberal arts colleges also qualify and are not listed), listed with counts on the site’s Patterns tab. The v0.4.2 figures were twenty-five and nineteen. Candidates were identified by link text and not opened, so this is a strong signal that these institutions publish deliberation routinely and a weaker signal about AI specifically; a few also expose AI-tagged links that proved on inspection not to be deliberative documents.

Watch items (bodies exist, no report posted): Colorado State (CSU System AI Joint Task Force, page unlinked, no document; the Provost’s AI Working Group’s draft guidance sits behind a SharePoint login and was the one unreachable verdict of the v0.5.0 triage); NC State (Provost’s AI Advisory Group, charged 2 April 2024; the charge is now a supplementary row); Washington (AI Governance Committee, charged 7 May 2026, final due 19 March 2027); Wyoming (AI Committee charge letters cite an unposted 2025 report; President’s Commission deliverable due June 2026); Notre Dame (Curriculum and Learning Task Force); Southern California (President’s AI Strategy Committee, roster only); Texas A&M (AI Ethics and Governance Working Group report on SharePoint); New Hampshire (two further senate AI documents on SharePoint); New Mexico (AI Steering Committee values statement of July 2026, now a supplementary row); Northern Arizona (June 2023 report referenced, unlinked); Idaho (white paper 404 after migration); Toledo; Vermont; New Mexico State (senate resolution of December 2025 requested a report by April 2026; the resolution is now a supplementary row); UMass Lowell; Penn State (three further senate reports, SharePoint-gated); UConn CLAS; Emory; Southern Mississippi (infrastructure recommendations forthcoming). R2: Appalachian State (Chancellor’s AI Task Force; the 2024 charter is now a supplementary row, the report is not posted); Villanova (final report on a personal OneDrive, now a supplementary row as gated); North Florida (AI Council and strategic plan pages, no report); Fordham (a second document behind sign-in). Loyola Marymount left this list in v1.5: the provost’s March 2026 letter carrying the AI Task Force recommendations is a headline row with access = summary, and the full document stays behind Box.

Not located under this protocol, not asserted absent: Harvard (provost report index back to 2004 contains only July 2023 guidelines that disclaim being policy; three university-wide working groups list charges and members and no reports), Columbia, Northwestern, Johns Hopkins, Georgetown, WashU, NYU, Carnegie Mellon, Michigan State, Purdue, UT Austin, Georgia Tech. UC Berkeley and Princeton, listed here in v1.0, were located in v1.1 and v1.2.

External audit (5 September 2026)

An independent structural audit of the live v1.0 site checked all inventory rows and census rows, attempted retrieval of every URL, inspected document metadata, and ran a targeted external search. Its findings are applied in v1.0.1 and the audit document is published alongside the data. Corrections: UNC-Chapel Hill date (August 2, 2024); Chicago 2025 recommendation removed as unsupported by the posted abstract; Dartmouth (single-author adviser report) and Ohio State Fisher (slide deck) moved to supplementary under the inclusion rule; Michigan 2023 and 2026, UNMC 2023 and UNLV 2024 added; Indiana marked systemwide; CU Denver attribution narrowed; Rochester revision dates recorded; genre, stage, access and scope fields introduced; census metric columns renamed for what the v0.2 crawler measured; crawler v0.3 written to capture off-domain documents and HTML report pages and to log per-URL outcomes; seed lists and raw crawl outputs published. Done in v1.1: the census re-run (with v0.4.1, which supersedes v0.3), a measured recall figure, and mechanical status assignment. Not yet done: a report–institution mapping for systemwide documents, and a Wayback snapshot of every URL.

v1.5.1: fixes from an external review (9 September 2026)

What was reviewed. An external pre-publication pass over the archived v1.5 release recomputed every distribution, ran roster and census diagnostics, and read the retained source text for three reports. Its numbers reproduced against the working files exactly. Four of its findings are matters for this release; the rest (charge relevance for the default code, an implementation-specificity reading of seriousness, clustered public/private comparisons, a human calibration queue) belong to the paper that will use the data and are not applied here.

Boston College moves to supplementary. Its document, Student Education and Empowerment in the Age of AI (June 2025), is a staff-authored account of a Student AI Advisors group convened by the Center for Digital Innovation in Learning and ITS, and its recommendations run to faculty practice. That fails part (a) of the inclusion test the same way CU Denver’s community of practice did. The review also found that the roster table had counted the report’s byline author as a staff seat. The row stays in the inventory as supplementary, linked and noted, and the benchmark holdout is not changed: the holdout tests whether the crawler finds a document, not whether the document is included, and altering it would break comparability across crawler versions. Headline set: 99 reports at 76 institutions, 84 at 63 R1 campuses.

Six chairs-only listings are no longer rosters. MIT, UC Davis, UCLA, Oklahoma, Virginia Tech and Sacramento State print their co-chairs and no membership list; the extraction recorded roster_present = false and the chairs, and the composition table had counted the chairs as a roster. roster-composition_v14.csv and coded_v14.csv now carry roster_type (full, leadership_only, none). The roster figures count the 71 full rosters: 1,484 seats, of which 521 print no title.

The not_stated default is qualified. Seven of the 47 rows are summaries, abstracts or an abridged text; see dimension 5.

The name scrub is tightened. A withheld name followed by its title and unit is withheld in spelling only. The publish script now removes the descriptor and cuts membership quotes at the first withheld name; see Privacy.

Every distribution was recomputed. Relative to v1.5: convening_authority other 6 to 5; student_role consulted_only 18 to 17; detector_stance caution 18 to 17; default_posture not_stated 48 to 47; syllabus_architecture silent 37 to 36; standing_body_rec no 23 to 22 (the dropped row was silent); procurement_stance build_own 7 to 6; ai_self_disclosure not_stated 74 to 73; convener_level campus 78 to 77; citations 1,330 edges and 871 nodes from 92 reports; scope teaching 87 to 86. The provost and senate standing-body shares (32 of 38, 5 of 12) and the detector shares are unchanged.

v1.5: crawler v0.5.0, the full re-crawl and the boundary rule (8 September 2026)

The crawler was fixed after a reader reported a miss. Boise State’s three AI Coordinating Council white papers had been fetched and rejected by v0.4.2. The diagnosis and the fix are under Coverage: regexes that could not match a plural, hosted documents that scored as if they had no file, and pages that named a body without lending that name to the documents they linked. On the frozen holdout the new crawler finds 44 of 55 institutions (80%) and 53 of 69 documents (77%), against 42 (76%) and 50 (72%) for v0.4.2 and 26 (47%) for v0.2. The cost is precision: the candidate pile is noisier and triage is the bottleneck.

All 372 institutions were re-crawled on 8 September. AI candidate links went from 720 to 1,752 and 66 institutions went from zero candidates to some. Eighty-six institutions without a report and with a candidate scoring 12 or more were triaged by hand under TRIAGE-RULES.md: include 3, supplementary 18 (three of them duplicating rows already present, so 15 new), exclude 64, unreachable 1 (Colorado State, where a Provost’s AI Working Group’s draft guidance sits behind a SharePoint login). The verdicts are in crawl_out_v050_census/. Every census row now carries v0.5.0 values for pages read, candidates and hosts; the status table under Coverage has a v1.5 column for each segment.

Five headline rows in, one out. Boise State enters with three white papers of the AI Coordinating Council (on assessment, on detection tools, on tutoring; April 2026, February 2026 and December 2025; genre guidelines_with_recommendations; hosted on Google Docs). Illinois Urbana-Champaign enters with “Generative AI Use Cases” (April 2024), an interim report of the Generative AI Solutions Hub working groups under the CIO. Loyola Marymount enters with the provost’s AI Task Force recommendations of March 2026, coded access = summary because the document is behind a Box login and the provost’s letter is the public text. CU Denver’s AI Faculty Liaison Community of Practice moves from headline to supplementary: a programme convened by three faculty fellows rather than a chartered body, and its recommendations run to units. The headline set goes from 96 reports at 74 institutions to 100 at 77 (R1 85, R2 14, LAC 1); R2 institutions with a report go from 10 to 12 of 139; supplementary rows from 9 to 25.

The boundary rule is now written down. Two parts, both required. First, a named deliberative body is the author; an office issuing guidance in its own name does not count, whatever the office consulted. Second, the recommendations run to the institution, not to the reader; a document that tells instructors what to do in their courses is guidance, not deliberation, however good. Three cases decided under it. Boise State’s white papers pass both tests: a named council is the author, and each paper closes with recommendations for institutional practice (the detection paper recommends against enterprise licensing of detectors, which is a recommendation to the university). UNC Chapel Hill’s ai.unc.edu guidance set fails the first: it is issued by the university in its own name, and the Chancellor’s Generative AI Working Group report of August 2024, already in the set, is the deliberative document. Penn State’s Faculty Affairs guidance on using AI to summarise student feedback for faculty review fails both, though it is consequential, and Iowa has a parallel document that fails the same way. Names do not decide: William & Mary’s “Artificial Intelligence Policy Initiative” passes, because it has a charge, a roster and a report to the Provost, and CU Denver’s “Community of Practice” does not. The rule is in TRIAGE-RULES.md and applied to every row carrying coding_basis = crawl_v0.5.0_triage.

Every distribution was recomputed. The changes, all traceable to the five rows in and one out: standing bodies yes_new_body 51, yes_new_office 8, yes_expand_existing 10, no 23, not_stated 8, with the 23 no rows adjudicated as 1 declines, 19 defer, 3 silent (the four new no rows all defer); convener_level campus 78, system 10, not evidenced 12; the citation network 1,334 edges and 874 nodes, with Microsoft (27) now one ahead of OpenAI (26) and Harvard 19 to Michigan 18 unchanged; rosters 1,508 seats across 78 reports; access full 90, summary 6, abstract 2, abridged 2. procurement_stance takes a sixth value, against_licensing, on one row (Boise State’s detection paper); see dimension 8. Each dimension below carries its v1.5.1 figures; the earlier release notes keep the figures they shipped with.

v1.4.4: campus and system separated (7 September 2026)

A new dimension, convener_level. convening_authority records the office title as printed, and a printed title does not say whether the charging office sits at a campus or above it. A chancellor runs the campus at UC Davis and the system in the Colorado State and California State systems. All 96 headline reports were read again for the charging office and coded campus, system, segment or not_evidenced, with the supporting quote, the extraction field the quote came from and a confidence flag published in convener_level.json. That file is now an input to code_v14.py, on the same terms as standing_body_adjudication.json, rather than a note about it. The distribution is campus 74, system 10, not_evidenced 12 and segment 0; of the ten system rows, four are high confidence and six low. See dimension 13 for the rule and the confidence split.

scope_level is coded independently and is now populated. It records whose people the recommendations bind, which is a different question from who chartered the body. Nine rows move from campus to system: Penn State, both Michigan reports, all three Minnesota reports, both Washington reports and Washington State. Indiana was already system, so the headline set goes from 95 campus and 1 system to 86 and 10. See dimension 13.

seriousness is displayed as a specificity score. The construct the column measures is specificity, not seriousness: it counts how many of three things a report does, name an owner, set a deadline, ask for money. The site now prints it as “specificity N of 3” and says on the page what the three are. The first audit recommended renaming the column to implementation_specificity. The name is kept deliberately, because renaming breaks a published column others may already have joined against; the gloss carries the correction instead. See dimension 9.

A chart for what the reports cover, built from the multi-valued scope field: teaching 83, governance 70, research 60, administration 49, curriculum 37 across the 96 headline reports. Reports carry more than one tag, so the counts do not sum to 96. Just over half the set, 49 of 96, reaches administrative and operational use, which is the point of the chart: the genre is not only about teaching. See dimension 12.

The joined map fields are documented, and one of them is withdrawn from analysis. A new section below describes map_quadrant, map_resilience and map_ai_exposure in crawl_coverage.csv and warns against the blended exposure column, which has no usable variance.

v1.4.3: the second audit (6 September 2026)

A second independent audit read the published v1.4.2 site and files against each other and against the documents behind them. Its corrections are applied to the data, to the site and to this file. Every count below was recomputed from coded_v14.csv, citations_v14.csv, roster-composition_v14.csv, the inventory CSV, crawl_coverage.csv and standing_body_adjudication.json.

Citation identity. Two guards, CITE_VENUE and CITE_GUARD, now run before the keyword map in norm_cite. A university’s name inside a string no longer makes that university the cited source: Harvard Business Review, Harvard referencing style and an author’s affiliation printed inside a press interview are no longer citations of Harvard, and the University of British Columbia and the University of Missouri-Columbia are no longer citations of Columbia. The network moves from 1,310 edges and 855 nodes to 1,312 and 858. Harvard falls from 21 to 19 against Michigan’s 18, one edge apart, and the claim that Harvard leads the universities is presented as a tie wherever it appears. See dimension 11.

Standing bodies. For the 19 reports coded no, the disposition is no longer computed from keywords. All 19 were read in full a second time and sorted by hand into declines, defers and silent, with the supporting passage and page in standing_body_adjudication.json, which is now a published input to code_v14.py. A new column, standing_body_disposition, carries the verdict: one report declines, fifteen defer and three are silent. The earlier claim that two reports decline in writing becomes one. CU Denver was re-read, and its “work within established policy” passage is about not duplicating policies rather than about refusing a body. Not one senate-convened report declines. standing_body_evidence moves from 61 stated and 35 inferred to 60 and 36. See dimension 7.

Procurement. license_enterprise_tool records a procurement model and not a vendor count. The display label becomes “license enterprise tools centrally” and the year chart becomes “share recommending central enterprise licensing”. Any reading of the counts as consolidation on a single tool is withdrawn, and the dimension now states which question it does not answer. The one report with no publication date is excluded from the by-year chart rather than plotted. Williams’s row was re-checked against its document and is flagged in dimension 8 as doubtful; the code is left as the script produces it. See dimension 8.

Recall. The R1 to R2 gap is no longer described as an upper bound: differential retrieval may affect the observed gap, and neither the size nor the direction of that effect is established here. The liberal arts finding no longer infers that those colleges rarely produce the genre. The claim that no R2 report was known before the R2 crawl is corrected. RIT’s report was harvested on 1 September 2026 and held in the supplementary set by an outdated R1-only rule, four days before the 5 September crawl, and the crawler read 146 pages there without surfacing it. One known positive is too few to estimate recall from, and that is the limitation the file now states.

Obsolete material cleared. The list of fields not yet populated no longer carries fields that are populated, committee composition among them. The claim that the Carnegie classification was not checked against a machine-readable file is removed: the denominators come from the 2025 Carnegie Research Activity Designations file, joined by UnitID.

v1.4.2: one promotion, two rule changes (6 September 2026)

Rochester Institute of Technology moves from supplementary to the headline set. It had been excluded with the note “Not R1 under the 2025 classification,” a rule v1.2 retired when the R2 segment was added; RIT is R2 in the 2025 Carnegie file and the exclusion should have lapsed then. The document is the RIT AI Task Force Final Report of April 2024, 81 pages, convened by the provost, with a standing AI Hub as its central recommendation. It was read in full and coded under the published rules. R2 headline reports go from 9 to 10, the headline set from 95 to 96, institutions with a report from 73 to 74, and the R2 census line from 9 found and 121 not surfaced to 10 and 120. Every distribution, count and crosstab in this file was recomputed from the published CSVs rather than incremented.

standing_body_evidence was rebuilt. Through v1.4.1 the field read quoted whenever any quote existed on the standing-body field, without testing whether the quote spoke to the question, which made the value meaningless. The rule and the value names both changed; the field no longer takes the value quoted. See dimension 7 for the rule and the new distribution.

The detector-by-convener figure was counting the wrong numerator. It counted detector_stance != 'silent' under a heading about taking a position, which folded in the nine reports that raise detection and take no position. It now counts only the four position values. See dimension 4 for the corrected shares.

Three descriptive corrections. The claim that every headline report was “read in full” is split into the two claims it was running together: what was read, and what the institution published. The voting_member label is defined against what the reports actually say. Roster totals are labelled as seats. UC Berkeley was removed from the not-found list at the foot of this file, where it had survived from v1.0 despite being located in v1.1.

v1.4.1: rosters withdrawn (6 September 2026)

The named rosters were withdrawn from publication a day after v1.4 shipped. rosters_v14.csv is replaced by roster-composition_v14.csv, and member names, titles and units are stripped from the published extraction and from every published table. Category and role remain, so code_v14.py reproduces the coded values from the published files. See Privacy for the reasoning. The Patterns charts were also rebuilt on fixed origins so that bar lengths are comparable within each family.

v1.4: the full-text pass (5 September 2026)

Counts in this entry are the ones v1.4 shipped, not today’s. All 95 headline documents in the set at the time were read, 93 end to end and two (UNC Charlotte, UW-Milwaukee) only as far as the public source exposes, which read_coverage records. Reading a document in full is not the same as reading a report in full: some of those documents are the summary, abridged version or abstract the institution published in place of the report, and access records which. Documents that refused automated retrieval were fetched through a browser or with a browser-impersonating client. Stage A produced one structured extraction per document (rosters, leadership, every recommendation, positions, citations); Stage B applied the published rules in code_v14.py to produce the ten coded dimensions, a 1,441-row roster table and a citation network whose size at that release cannot now be reconciled: the v1.4 page reported 785 edges and the v1.4.1 page 1,267, and neither reproduces from the v1.4.2 table. The dimensions were chosen from what varied in the extraction, not before it; four earlier candidate dimensions were dropped as near-constants. Reading the documents also corrected the convening authority on 21 rows. Nothing in the census or the inclusion set changed; v1.4 is a coding release. RIT’s promotion in v1.4.2 and the citation-identity guards in v1.4.3 account for the difference between these figures and the current ones.

v1.3: selective liberal arts colleges (5 September 2026)

A selection, not a Carnegie class, and stated so it can be argued with. From the 216 institutions in the 2021 Baccalaureate: Arts & Sciences category (C21BASIC = 21): at least 300 undergraduates, endowment per undergraduate of at least $150,000 (College Scorecard ENDOWEND / UGDS), and an admit rate of 40% or lower (ADM_RATE). Forty-six colleges. The joint rule was chosen over endowment alone because endowment alone admits institutions that are rich per student for reasons unrelated to selectivity (Soka, Principia, VMI, Earlham), and over selectivity alone because the argument the segment is meant to test is about who can afford instructor discretion. The $150,000 line rather than $200,000 brings in Spelman, Barnard, Pitzer, Skidmore and Gettysburg at the cost of no additional noise. Institutions just outside the line, in order of how narrowly they miss the admit-rate cutoff and setting Soka aside: Hampden-Sydney, Dickinson, Furman, Southwestern, Millsaps, Union, Occidental, Austin College, St. Olaf, Rhodes. The selection variables are in lac_all.json.

Result. Crawled with v0.4.2; three sites refused the crawler (Colby, Holy Cross, Williams) and were checked by hand, at most six page loads each, disclosed in the census notes. Three crawler candidates and two manual finds were triaged. One headline report: Williams, Final Report from the 2025-2026 Ad Hoc Committee on AI and Academics (July 2026, 16 pages; twelve members including two students; convener not named). Supplementary: Amherst’s Fall 2023 task force report (behind community login; charge and membership public) and Colorado College’s two-page Generative AI Institutional Philosophy (principles, not findings). Watch: Holy Cross (an Institutional Review of AI Task Force reported in the college magazine, April 2025; interim ethical-use statement intranet-only), Colby (a General Counsel memo and an unattributed working paper), Lafayette (2023 working-group guide behind sign-in), Amherst (College-wide AI Working Group chartered 2025-27). Spelman is the fifteenth HBCU in the census; none of the fifteen has a findable report.

Link check at v1.3. Every inventory URL was re-fetched. Eleven did not return a clean 200 from this environment; nine are bot-blocks that open in a browser (Minnesota’s repository, Buffalo, UNMC, Williams, Amherst, UMKC) or gated originals recorded as such. Two changed the data: UNLV’s May 2024 Faculty Senate report, added from the external audit as a public Google Drive file, now requires sign-in (confirmed in a browser) and moves to supplementary as gated; the Princeton and Wake Forest rows now link the public summary (a dean’s memo; an adapted guidelines page) rather than the gated original, which is kept in the note.

Reading the negative space. A single report across 46 colleges is not evidence that these colleges have not deliberated about AI, and it does not establish how often they produce this genre either. Recall on liberal arts sites is unmeasured, three of the 46 refused the crawler and were checked by hand at no more than six page loads each, and one findable report is a count rather than a rate. The available reading is that small faculties govern by faculty meeting, Committee on Educational Policy and honor council, and that the output is a vote, a handbook amendment or a teaching-center page, none of which passes the inclusion rule. That reading is a hypothesis the census cannot test. A follow-on layer coding institutional AI posture from honor-code and academic-integrity language would measure what this segment actually decided; it would also be a different project, and the codebook does not claim it.

v1.2: Research 2 (5 September 2026)

All 139 R2 institutions in the Carnegie 2025 file were crawled twice with the same crawler as R1: a first pass under v0.4.1, which exhausted its queue after the seeds at 52 sites, and a second under v0.4.2 after the path-following fix, which read a median of 145 pages per site. Thirty-one candidates were triaged under the published rules: include 8, supplementary 4, exclude 17, unreachable 2. One supplementary verdict (East Tennessee State’s task-force-authored guidance with recommendations) was promoted under the boundary-genre rule; James Madison’s interim update is a supplementary row; the Appalachian State charter and Villanova’s gated report are watch items. The R1 set was re-crawled under v0.4.2 for consistency, which surfaced Princeton (summary of a gated report) and UMKC. Rows added in v1.2 carry coding_basis = crawler_v0.4.2_triage.

v1.1 census re-run (5 September 2026)

All 187 institutions were re-crawled with crawler v0.4.1. One hundred seventy-nine candidates carrying all three signals and not already in the inventory were opened and triaged by Claude under the rules in TRIAGE-RULES.md (published with the verdicts in triage_v0.4.1.zip): include 8, supplementary 16, duplicate 42, unreachable 12, exclude 101. Of the 16 supplementary verdicts, seven committee-authored guidelines with recommendations were promoted to the headline set under the boundary-genre rule already applied to Kentucky, LSU and Rochester; six charge letters became watch items rather than rows; one college-level guideline and one companion whitepaper are supplementary rows; one computing-institute summary was dropped as not about AI. Fifteen headline rows were added at ten institutions, seven of them previously recorded as having no findable report. Rows added in v1.1 carry coding_basis = crawler_v0.4.1_triage and date_accessed = 2026-09-05; their dimension codes are uncoded pending the full-text pass.

Privacy: what is withheld and why

Every member name in these rosters is already public, printed in a report its own institution posted. Compiling them into one searchable index is a different act. It produces a list of people identified by name, title, unit and position on an AI committee, which is precisely the artifact a spammer or a harasser would want and which no single report supplies. The research question is answered by composition, not by identity, so composition is what is published: aggregate counts, and per-report counts by category and role in roster-composition_v14.csv. Member names, titles and units are also stripped from the per-document extraction published under Data and Downloads; the category and role fields remain, so code_v14.py still reproduces every coded value from the published files. Individual names remain readable in each linked report, in the context the institution chose. Researchers who need the named rosters can ask.

The same reasoning applies to the coded table, where individual officers named in a transmittal or a quotation are reduced to their office. It does not apply to citations: a report cited by its authors’ names is a bibliographic reference, and those stay.

From v1.5.1 the scrub also removes the title and unit that follow a withheld name in evidence text. A name withheld in spelling only, with “Chair of the Academic Senate” or an endowed chair left standing beside the placeholder, identifies the person as surely as the name did, once the institution and year are known; the external review of v1.5 counted about eighty such strings. Membership quotes are cut at the first withheld name. Offices are kept: a provost, president, chancellor or dean named as the convener or the addressee is a public officer acting in office, and the office is what the coded fields record. publish_v142.py asserts that no withheld name is followed by a descriptor before it writes the upload set.

Reliability: stated plainly

Author spot-check (v1.3 and v1.4). On 5 September 2026 the author re-checked 14 self-drawn inventory rows against the linked documents (institution, body, title, date, link) and concurred with all 14, and reviewed the v1.4 coded values against the documents they were drawn from and concurred with those as well. One person; recorded because earlier versions said the check was pending.

There is no independent second coder and there will not be one. This is a one-person project. Releases through v1.3 carried only fields verifiable from the document or its landing page. The v1.4 coding was performed by Claude in two separated stages (extraction, then rules) described under Coded dimensions; the author reviewed the codes against their documents and concurred (see Author spot-check). No inter-rater statistic will be reported, because none can honestly be computed from one coder plus a check. Anyone who needs coded variables at a reliability standard should treat v1.4 as a starting point and re-code from the linked documents: the full text of every report, the extraction, and the coding script are all published, so a second coder can start where this one did.

Fields

institution · state · control (public/private) · carnegie_2025 (R1/R2) · body_name · convening_authority · report_title · date_published · date_accessed · url · url_http_status · scope · scope_level · convener_level · key_recommendation · plus the coded dimensions below.

IPEDS UnitID is populated for every census row and every headline inventory row, and the Carnegie class is read from the 2025 Research Activity Designations file and joined on that UnitID, not name-matched from a published list. pages is populated in coded_v14.csv for the 77 of 99 headline documents that expose a page count, with an observed range of 2 to 128. Committee composition is published per report in roster-composition_v14.csv and summarized under dimension 3.

Joined map fields

crawl_coverage.csv publishes three columns joined by IPEDS UnitID from Mapping the Structural Divide, a separate project: map_quadrant, map_resilience and map_ai_exposure. They are carried so that a reader can ask whether the institutions producing this genre differ structurally from those that do not. Coverage: 344 of the 372 census rows carry all three values, two more (Claremont Graduate University, Teachers College at Columbia University) carry a resilience score and no exposure value, and 26 do not join. Quadrant distribution across the 344: High Capacity 243, Market Misaligned 36, High Stress 36, Structurally Exposed 29.

Do not analyse map_ai_exposure. The column is that project’s AI_EXPOSURE_BLENDED, a composite whose standard deviation across all 1,609 institutions in the source dataset is about 0.014, on a range running from 0.39 to 0.51. Every institution scores near 0.46, so the column has no usable variance and any gap between two campuses in it is smaller than the third decimal place. Within the 344 census rows the standard deviation is smaller still, about 0.008. A correlation run against this column will return noise and should not be reported. The component that does vary is L_AIEXP, the inverted entry-level task-exposure score, with a standard deviation of about 0.29 across the same 1,609 institutions. Anyone who wants an exposure measure should take L_AIEXP from the source dataset rather than the blended column published here.

The quadrants are unaffected. map_quadrant is a median split on two axes, resilience and labor alignment, and the blended column is an input to neither. L_AIEXP feeds the labor alignment axis. map_resilience carries a standard deviation of about 0.14 across the census rows and a range from 0.31 to 0.92, so it behaves as a variable.

Coded dimensions: released in v1.4, from the full text

Every document in the headline set was read: 96 of the 99 end to end, and three (UNC Charlotte, UW-Milwaukee, Loyola Marymount) only as far as the public source exposes. That is a claim about documents, not about reports. Ten of the 99 documents are the summary, abridged version or abstract the institution published in place of the report, so for those ten the codes describe a public stand-in and not the underlying report; access marks them and read_coverage records how much of each document was read. The pass had two stages, deliberately separated so the coding scheme could not be fitted to the first fifty documents read:

Stage A, extraction. Claude agents read each document (96 end to end; three only as far as the public source exposes) and recorded facts and verbatim evidence into a fixed schema (EXTRACTION-SCHEMA.md): the charge and who signed it, the membership roster as printed, each member’s category and role (names, titles and units are withheld from the published copy; see Privacy), who the body reports to, who owns implementation, every recommendation with its owner, deadline and resource ask, positions on detectors, syllabus requirements and defaults, tools named, policies named for amendment, works and institutions cited, and whether the report says it used AI to write itself. Every judgment field carries a quote and a page. Where the document was ambiguous, the extractor recorded both readings in extractor_notes rather than choosing.

Stage B, coding. The dimensions below were chosen after the extraction existed, from what actually varied in it, and each is a deterministic function of extraction fields computed by code_v14.py. Every dimension but one is computed without a per-document judgment call: where the extractor flagged an alternative reading, the published rule decides. The exception is the disposition of the 22 reports coded no on standing bodies, which v1.4.3 moved out of a keyword test and into a hand adjudication published as standing_body_adjudication.json; the script reads that file rather than deciding for itself, and dimension 7 gives the reasoning. Running the script over the published extraction and that file reproduces coded_v14.csv, citations_v14.csv and roster-composition_v14.csv exactly. It also writes rosters_v14.csv, which carries the member names and is not published; run against the published extraction, that file comes out with its name fields empty.

Four dimensions that had been in the scheme since v1.0 were dropped as near-constants and are not coded: faculty training (80 of 100 “recommended” and 4 “mandatory”), student AI literacy (67 “recommended” and 15 a curricular requirement), assessment redesign (57 yes, 43 silent), and the problems a report names, which are close to universal (privacy and data 87, pedagogical opportunity 80, integrity 79, equity and access 77). They remain in the extraction for anyone who wants them; the counts just given come from the extraction JSONs, not from a published CSV column.

1. convening_authority: who chartered the body

provost 38 · joint 16 · faculty_senate 12 · unknown 10 · president 7 · other 5 · chancellor 4 · research_office 3 · cio_it 3 · dean 1

Rule: the office named in the document’s own charge or transmittal. Where the document does not name one, the value falls back to the landing page or transmittal used in v1.3 and convener_source records which (document 84, landing_page 6, none 10). The v1.3 value is retained in convening_authority_landing. Of the 95 rows that carry a landing-page value, reading the documents changed 20, left 67 unchanged, and eight more differ only in vocabulary; the five rows added in v1.5 have no landing-page value to compare. Boise State’s three white papers name no charging office and account for three of the ten unknown. The changes are not noise: several bodies that a landing page credited to a provost are jointly chartered in their own charge letters (Michigan by the Provost and the VPIT-CIO, Virginia Tech by the Provost and the COO, Penn State by the Provost and the Faculty Senate Chair), and two that a page credited to a president are signed by a president and a provost.

other covers bodies chartered by something that is none of the above: a teaching center (Boston College), a self-organized faculty group (Penn’s Price Lab), a Chief AI Officer (Oklahoma), a summit (Texas A&M). convener_as_printed carries the string in every case.

The title alone does not say whether the office sits at a campus or above one, and this dimension does not try to. That question is coded separately in convener_level; see dimension 13.

2. joint_composition: what “joint” means, when it applies

admin_and_senate 6 · provost_and_research 3 · provost_and_cio 3 · provost_and_operations 2 · president_and_provost 2

Kept as a second column rather than splitting convening_authority, because splitting leaves cells of two and three and makes every crosstab unreadable. Assigned by keyword from the printed convener string, first match wins, in the order listed in code_v14.py.

3. student_role: where students sit relative to the body

not_evidenced 31 · absent 27 · voting_member 24 · consulted_only 17

Rule, in order: a student named on the roster as a member or chair is voting_member; a body whose own name says it is a student body is voting_member, unless the report states a narrower role for students, in which case the stated role governs; if the roster is untyped (half or more of its entries print no title) or partial (shorter than the membership total the report states, or two names or fewer, which is a signature block rather than a membership list) the role stated in the report’s prose governs; otherwise a complete typed roster with no student is absent. voting_member means the report seats a student as a member of the body; it does not assert that the student holds a vote, which most of these reports never address, and the published display label for the value is “seated as a member”. The not_evidenced cases are documents that leave the question open: 20 of the 31 print no roster, and the other eleven print one that is untyped or a signature block while the prose says nothing about students. They are a finding about the documents, not a measurement of the bodies.

Across the 71 reports that print a roster, 1,484 seats: 429 faculty, 292 administrators, 207 staff, 27 students, 8 external, and 521 whose title is not printed. Six further reports print only their co-chairs; from v1.5.1 the composition table marks those leadership_only in a roster_type column, and they are excluded from every roster figure. These are seats, not deduplicated people: a member who sits on two reports’ bodies is counted twice, and no attempt is made to resolve names across documents. Per-report composition is in roster-composition_v14.csv. The names themselves are not published; see Privacy below. Six of the 24 voting_member codes rest on a report that says students served without naming which member is the student, which students_n = 0 marks in the coded table.

4. detector_stance: position on AI detection tools

silent 42 · recommend_against 18 · caution 17 · one_input_among_several 12 · mentions_no_position 9 · recommend_use 1

Rule: the extractor’s position field, with one split added after a continuity check found nine documents whose position was recorded as “not mentioned” while the extraction also carried a detector sentence. Those are coded mentions_no_position: detection comes up, the report takes no position on using it. silent now means the document does not raise detection at all.

Forty-two reports never raise detection, and nine raise it without taking a position. Of the forty-nine that take one, eighteen recommend against detectors, eighteen caution, twelve treat detector output as one input among several, and exactly one (Maryland) recommends piloting one.

Rule change, v1.4.2: what counts as taking a position. The detector-by-convener figure counted detector_stance != 'silent' under a heading about taking a position, which folded the nine mentions_no_position reports in with the forty-nine that state one and overstated every share. The numerator is now the four position values (recommend_against, caution, one_input_among_several, recommend_use) and nothing else. Shares by convener at v1.5.1: provost 23 of 38, joint 8 of 16, faculty senate 6 of 12, president 3 of 7, chancellor 3 of 4. mentions_no_position sits with silent on this count, because a report that raises detection and declines to say anything about it has not taken a position on using it.

5. default_posture: the institution’s default rule for student use

not_stated 47 · instructor_sets 42 · prohibit_unless_permitted 9 · permit_unless_prohibited 1

The most consequential negative result in the set. Only ten of 99 reports state an institution-wide default at all, and only one (Michigan 2023) defaults to permission. Two qualifications on the 47 not_stated rows, both from the v1.5.1 review: seven of them are summaries, abstracts or an abridged text, so their silence is silence in what the public source exposes and not in the report; and a body charged with research, procurement or infrastructure had no reason to settle student use, so the code records that the document does not state a default, not that the body declined to set one. Whether the question fell within the charge is not yet coded. “Banning AI” is not what this genre does; pushing the decision to the instructor is.

6. syllabus_architecture: how course-level rules are structured

silent 36 · recommended_not_required 30 · every_syllabus_must_state 22 · tiered_menu_offered 10 · instructor_discretion 1

The operational form of the posture question: 22 reports would require a statement in every syllabus.

7. standing_body_rec: does the report recommend permanent machinery

yes_new_body 51 · no 22 · yes_expand_existing 10 · not_stated 8 · yes_new_office 8

What a no records. A no records that no proposal for a standing body appears in the report. It does not record that the report refused one. In nearly every case the document simply does not raise permanent machinery, and reading a decision into that silence is reading something the document does not say. The companion fields below are what separate the two.

standing_body_disposition, new in v1.4.3: the no rows, read again and sorted by hand. Through v1.4.2 the disposition of a no was decided by a keyword test, SB_DECLINE, which checked for the presence of words rather than for what a sentence claimed. It read CU Denver’s “designed to work within established policy” as a refusal, when that passage is about not duplicating policies and says nothing about creating a body. The test is retired for these rows. All 19 reports coded no at v1.4.3 were read in full a second time and decided one at a time, and the four no rows added in v1.5 (Boise State’s three papers and Illinois) were adjudicated the same way on entry, with the supporting passage and its page recorded in standing_body_adjudication.json, which is published beside the script and is now an input to code_v14.py rather than a note about it. The verdict is carried in a new column, standing_body_disposition, which is empty for every other value of standing_body_rec. The three verdicts are definitions, not impressions:

The adjudication is single-coder and done by hand, on the same terms as every other judgment in this file (see Reliability). Its evidence is published so that any verdict can be argued with against the document it came from. Not one senate-convened report declines: of the seven senate no values, six defer and one (San Diego State) is silent.

Rule change carried forward from v1.4.2: standing_body_evidence. Through v1.4.1 this field was set to quoted whenever any quote existed on the standing-body field, without testing whether the quote spoke to the question, which made the field meaningless. The values are stated and inferred, not quoted and inferred. For the three yes_ values and for not_stated, the test is SB_STRUCT in code_v14.py: a quote returns stated only if it names a standing structure, matching committee, council, office, board, hub, centre or center, institute, body or bodies, task force, working group, and the modifiers standing, ongoing, permanent, oversight, steering, governance and advisory. A missing quote, or a quote that names no structure, returns inferred. For the 22 rows coded no the value is no longer computed from the quote at all: it follows the hand adjudication, where declines maps to stated and both defers and silent map to inferred.

Distribution at v1.5.1: stated 61 · inferred 38. Among the 22 reports coded no, 1 is stated and 21 are inferred. At v1.4.3, when the set held 96 reports and 19 no rows, the old rule would have returned quoted for 86 of the 96, including 15 of the 19, because 86 rows carried some quote on the field; the difference is quotes that do not speak to the question.

Crossed with the convener, this is the governance finding, with the qualification above attached to every no. Counting any recommendation for standing machinery (new body, new office, or an expanded existing one), provost-chartered bodies recommend one in 32 of 38 cases and joint bodies in 14 of 16; counting only brand-new bodies, 22 and 13. Faculty senate committees recommend some form in 5 of 12, and the remaining 7 of 12 are coded no. All seven of those are inferred, and six of them defer to structures that already exist and one as silent: not one declines standing machinery in words. The one stated no in the whole set is Sacramento State’s, and that body was chartered by a provost. Administrations institutionalize. Senates deliberate and, on the evidence of the documents, stop short of proposing machinery, which is a weaker claim than refusing it and is the only one the data supports.

8. procurement_stance: what the report says to buy

license_enterprise_tool 32 · silent 31 · agnostic_multiple 20 · defer 9 · build_own 6 · against_licensing 1

What license_enterprise_tool records, and what it does not (v1.4.3). The value records that the report asks the institution to license enterprise access centrally, rather than leaving individuals to buy their own. It says nothing about how many vendors. RIT asks for the enterprise versions of ChatGPT and Copilot; USC names ChatGPT, Gemini and Copilot; Michigan Tech writes that its working group is not opposed to subscriptions to other comparable tools should a better fit emerge; UMKC asks the university to select and standardize a primary and a secondary general-use tool and to define a buy-versus-build strategy. The display label is “license enterprise tools centrally” and the year chart is “share recommending central enterprise licensing”. Procurement model and vendor concentration are separate questions, and this dimension separates neither: a report naming three products and a report naming one sit in the same category, and nothing here counts how many tools a campus ends up running. Any reading of these counts as evidence that campuses are consolidating on a single tool is withdrawn.

A sixth value, against_licensing, new in v1.5. Boise State’s white paper on detection tools finds no value added sufficient to justify enterprise-level adoption of AI detection tools. That is a procurement position, and it is neither defer nor agnostic_multiple, so the script now emits against_licensing when the extractor’s position field reads against. One row carries it. It is a position on one class of tool, not on enterprise licensing in general, and the same council’s other two papers are silent and agnostic_multiple.

The by-year chart. Its denominator is the reports that state a procurement position, which is every value other than silent, and the one report with no publication date (UW-Milwaukee’s Responsible AI Workgroup report) is excluded rather than plotted in an “unknown” column: 3 of 5 in 2023, 12 of 22 in 2024, 10 of 22 in 2025, 6 of 19 in 2026. Yearly counts are small and the 2023 bar rests on five reports.

One row is doubtful and is left as coded. Williams is coded license_enterprise_tool on a sentence from its peer survey, which reports that peers are moving “away from licensing individual commercial platforms” and toward internal portals. That sentence describes other colleges, and it points away from central licensing rather than toward it. The report’s own recommendations ask for a leadership structure, institutional principles, shared norms, faculty grants and student engagement, and none of them asks the College to license anything centrally. The code is published as code_v14.py produces it and is flagged here rather than overridden by hand.

9. seriousness: a specificity score, 0 to 3, one point each

+1 if any recommendation names an owning unit or role; +1 if any carries a deadline; +1 if the report prints a dollar figure or an FTE ask. 3 45 · 1 27 · 2 21 · 0 7. The component counts (owners_named_n, deadlines_n, dollar_figures_n, fte_asks_n) are in the coded CSV, so the index can be rebuilt differently by anyone who disagrees with the weighting.

The construct is specificity, not seriousness (v1.4.4). The score counts how many of three things a report does, name an owner, set a deadline, ask for money, and nothing else. It says nothing about how seriously a body took its charge, how well argued the report is or whether anything was implemented. A three-page memo that routes one recommendation to a named office by a named date outscores a hundred-page deliberation that does neither. The site prints the value as “specificity N of 3” and states the three components beside it. The first audit recommended renaming the column to implementation_specificity; the name is kept deliberately, because renaming breaks a published column that others may already have joined against, and the gloss carries the correction instead.

10. ai_self_disclosure: does the report say it used AI to write itself

not_stated 73 · yes 25 · no 1

Twenty-five reports disclose using generative AI in their own drafting (Michigan, Missouri, Minnesota, UC San Diego, William & Mary, Rhode Island, Tarleton, all three Boise State papers and others); one (Cornell’s research task force) states that it did not.

11. Citation network: citations_v14.csv

The network runs 1,330 edges from 92 of 99 reports to 871 normalized nodes. An extraction entry naming several organizations at once (“NSF, NIH Bridge2AI, USDA NIFA”) is split before normalizing, and a published keyword map collapses spellings; the printed string stays in cited_raw so the map can be redone. An edge is a report citing a source, so a report that names the same source twice contributes one edge.

Two guards run before the keyword map, new in v1.4.3. A university’s name inside a string does not make the university the cited source. CITE_VENUE routes a string to the venue that carries the name, so Harvard Business Review is Harvard Business Review, “APA, Chicago, Harvard style guides” is a citation style rather than a university, and an author’s affiliation printed inside a press interview does not become a citation of that author’s institution. CITE_GUARD protects longer institution names that contain a shorter one, so the University of British Columbia and the University of Missouri-Columbia no longer count as citations of Columbia. Totals move from 1,310 edges and 855 nodes to 1,312 and 858: the guards strip misattributed edges from the university nodes and add nodes of their own, and two reports that cited both a university and a venue carrying its name now contribute two edges where they used to collapse into one.

Most cited at v1.5: Microsoft 27, OpenAI 26, University of California (system or campus) 23, National Science Foundation 21, Harvard University 19, University of Michigan 18, Stanford University 17, then EDUCAUSE and Cornell University at 15, then three at 14 (National Institutes of Health, White House executive orders, Google). At v1.4.4 the two vendors were tied at 26; Boise State’s papers moved Microsoft one ahead, and the same caution about one-edge margins applies to that lead.

The two most cited sources are vendors, one edge apart. Harvard at 19 and Michigan at 18 are the two most cited individual universities, and one edge separates them: read that as a tie, not as a ranking. The margin will not bear weight, because a single sentence in a single report moves it, and because the guards just described moved Harvard by two on their own. Harvard has no report in this inventory: what circulates is its July 2023 guidance, which disclaims being policy. Michigan’s 2023 report is the single document most often cited by name. The University of California ranks above both only because its campuses and system office pool into one node.

What the nodes still mix. Normalization makes spellings comparable; it does not make sources comparable. One node list holds institutions, membership organizations, government agencies, publishers and named works side by side, and no field separates them. A single benchmarking sentence that lists many peers still generates one edge per peer, so a report that names ten peer institutions in passing contributes ten edges of the same weight as ten substantive citations. Counts here measure how often a name appears across reports, and nothing more.

12. scope: what the reports cover

Semicolon list: teaching · research · administration · curriculum · governance. Separates “we wrote a syllabus policy” from “we reorganized.” Assigned in v1.0 from the document’s own framing.

Tag counts across the 99 headline reports at v1.5.1: teaching 86 · governance 74 · research 62 · administration 51 · curriculum 37. The field is multi-valued and reports carry more than one tag, 311 tags in all and 3.1 per report on average, so the counts do not sum to 100 and no percentage should be computed against that denominator without saying so. Teaching is close to universal and curriculum is not, which is the distance between advising instructors on the courses they already teach and changing what is taught. Just over half the set, 51 of 99, reaches administrative and operational use, so the genre is not only about teaching: a substantial share of these bodies were asked about procurement, records, admissions and staff work as well. The tags are the document’s own framing rather than a reading of its recommendations, so a report that mentions research in its charge and never returns to it still carries the tag.

13. convener_level and scope_level: campus, or above it

convener_level: campus 78 · system 10 · not_evidenced 12 · segment 0
scope_level: campus 90 · system 10

Two questions the printed office title runs together. convening_authority records the title as printed, and a title is not a level: a chancellor runs the campus at UC Davis and the system in the Colorado State and California State systems, a president runs a single campus at one institution and a multi-campus system at the next. All headline reports were read for the charging office and coded, with the supporting quote, the extraction field it came from and a confidence flag published in convener_level.json, which code_v14.py reads as an input.

The rule for convener_level. A row is system in either of two cases. The first is that the named charging office is itself a system office, a system president or chancellor or a systemwide senate, and that is the high-confidence case, four of the ten. The second is that the document evidences that the charging office chartered a body whose remit spans more than one campus, and that is the low-confidence case, six of the ten. Any other named office is campus. A campus counts toward system only when it has its own chief executive under a common president; multi-campus language on its own does not qualify, which is why Cornell, Oklahoma, Northeastern and UMass Amherst are campus despite writing about more than one campus. segment is defined for a charging office above a system, a coordinating board or a state segment, and no report in the headline set has one, so the value is published empty. not_evidenced covers the 12 documents that name no charging office at all; those rows carry no confidence flag.

Confidence overall: high 70, low 18, and empty for the 12 not_evidenced rows. Among the 78 campus rows, 66 are high and 12 low. The evidence field records where each verdict came from: charge.quote 55, charge.convener_as_printed 23, fulltext 4, document 4, leadership.reports_to 1, landing_page 1, and none for the 12. The document and landing_page labels arrived with the v1.5 rows, whose verdicts were recorded at triage rather than from the Stage A extraction fields; the quote is published either way. Every quote is published, so any verdict can be argued with against the sentence it rests on. The site labels the three values “chartered at the campus”, “chartered above the campus” and “no charging office named”.

scope_level is a different question. It records whose people the recommendations bind rather than who chartered the body, and the two can disagree in either direction: a system office can charter a body that reports on one campus, and a campus body can write something a system later adopts. Coding them separately is what makes the disagreement visible. In the published data the ten reports with system scope are the same ten with a system convener, and the 12 rows with no evidenced convener all bind a single campus, so the two columns do not in fact part company anywhere in this set; that is a finding about these 99 documents and not a property of the coding. Nine rows moved from campus to system in v1.4.4: Penn State, both Michigan reports, all three Minnesota reports, both Washington reports and Washington State. Confidence on scope_level is high for 92 rows and low for seven. A third value, college_unit, appears three times in the inventory CSV and on no headline row: all three are supplementary documents written for a single college (Ohio State Fisher, Alabama Graduate School, and one v1.5 triage row).

⚠️ Coding validity: read before using

silent and not_stated now mean what they say. Through v1.3 these codes meant “not present in the harvest summary.” From v1.4 they mean the document was read end to end and does not take the position. The two are not comparable; do not chain v1.3 and v1.4 counts.

Three limits remain. First, single coder: the extraction and the coding were both done by Claude, under published instructions, with no independent second reader (see Reliability). Second, the unit is the report, so an institution whose provost’s task force is silent on detectors may still have a detector policy elsewhere; these codes describe documents, not institutions. Third, ten of the 99 documents in the headline set are the summary, abridged version or abstract the institution published in place of the report, and their codes reflect what the public document says, not what the gated full report may say; access marks them and stage records where the report sits in its own sequence.

Fields not yet populated

Recommended by the prior-art scan, deliberately left empty rather than guessed. Fields listed here through v1.4.2 that are now populated have been removed from the list: IPEDS UnitID, Carnegie class, page count and committee composition are all populated and are described under Fields and under the coded dimensions. - Archived snapshot URL + content hash. Urgent: UNC Charlotte’s URL contains backup-aihub, and UF’s report moved from .pdf to .docx, breaking the URL cited in the FAccT paper. Link rot is already happening. - Charge date → enables charge-to-report interval (College of Southern Nevada’s was 28 months). - Word count. Page count is populated and is the cheaper proxy for institutional investment; word count is not.

Not-found list (weak negatives: see the discoverability warning above)

I initially recorded Colorado State University as a verified absence, having checked ai.colostate.edu and found a mission/vision essay, a centers list, CSU-GPT, resources and news, but no report. That was wrong. A CSU report exists. It is simply not linked from the institution’s own AI hub.

This invalidates the method behind every negative in this file. The harvest searched AI hubs, provost pages and faculty-senate sites; a report that lives outside those, on a committee page, a system site, a strategic-planning microsite or an unlinked PDF, will not be found. An institution’s AI hub is not an index of its AI deliberation. Several institutions listed below as “expected and not found” almost certainly have reports that this method missed.

Every negative in this section should be read as not linked from the surfaces I checked, which is a much weaker claim than absence. Closing this properly requires per-institution site search plus a direct query to the provost’s office, not crawling.

Report these as “searched and not found,” not as silence.

Provenance

Harvested 2026-09-01 by six parallel searches (R1 public, private, regional/community college, system/association, international, prior-art). Every URL was opened by a harvesting or crawling agent, and each of the 128 inventory rows carries a status in url_http_status, recorded at the v1.3 link check for rows present then and at triage for the 23 rows added in v1.5. Four return 403 to datacenter IPs and open normally in a browser (UMN Conservancy ×2, Buffalo ×2).

Non-R1 segments (regional publics, community colleges, state systems, associations, accreditors, international) were harvested and parked in supplementary.html (linked from Data & Downloads). That file contains the accreditor findings (MSCHE’s binding AI policy effective 2025-07-01, WSCUC, C-RAC), which are directly relevant to accreditation work.