Compliance notice: Every prediction here is generated by a model, and a score is real only where its source is named. Nothing on this site takes money or accepts a wager.

Where the scores come from

A score on this site is real only where its source is named beside it. Everything else — every prediction, and every result with no credit — is generated by the model. There are three named sources: a credited publication, the stands, and coaches. These are they, and what each one is allowed to be.

Collected under its robots policy

Cal-Hi Sports

More than 1,700 credited games

Seasons since 2013, mostly California. Read under the site's own robots policy, and each score is credited and linked to the page it came from.

calhisports.com
Reported at the game

People in the stands

Settled by agreement

Anyone at the ground can report a score. It shows once two strangers agree, and a final that three separate voices agree on with enough standing becomes the game result, credited “via the stands”.

Report a score
Sent in, then reviewed

Coaches and contributors

0 games across 0 states

A coach sends a whole season; a reader sends one result with a link to where they saw it. A person reads every row before it counts, and the contributor is the credit.

Send a season

Every source assessed

The registry the collector runs against, served live. Refusals are listed with their reasons because a source that is missing looks like an oversight and one that is here with its reason looks like a decision.

  • contributions

    May collectchecked 2026-09-21

    Accepted submissions from verified people, or from anyone once a reviewer has read the row. The facts are theirs to report.

  • supplied-files

    May collectchecked 2026-09-21

    CSV files the operator dropped in PGSE_RESULTS_INBOX. Their entitlement to the data, not ours, and nothing is fetched.

  • scorestream

    Not yet askedchecked 2026-09-21

    A public API carrying 10,000-15,000 high school games a week, crowd-sourced and syndicated to the Associated Press. Access is granted per partner rather than scraped: partner@scorestream.com issues a developer key. Nobody has asked yet — this is the one entry here that is a conversation rather than a refusal.

  • texas-public-information-act

    Not yet askedchecked 2026-09-21

    Not a feed and not a favour. UIL was created by and is administered at UT Austin, which makes it a governmental body under Tex. Gov't Code ch. 552, and every Texas school district is one without argument. A terms-of-service clause is a contract; a records statute is a right, and the custodian does not get to decline. See docs/records_requests.md for the request to send.

  • khsaa-scoreboard

    Refusedchecked 2026-09-21

    The find of the survey and it still does not open. KHSAA publishes the Riherds scoreboard, 1998-2026, and khsaa.org/robots.txt permits this crawler — the first game-level source in the country that does. The scoreboard itself is three vendors deep (khsaa.org -> khsaa.arbiterwebsites.com -> rschooltoday.com) and Arbiter's terms forbid users to "copy or download any material or information from the Arbiter Services without the prior written consent of Arbiter". Same shape as MaxPreps: robots permits the crawl, the terms forbid the use. The records themselves are a different question — see khsaa-open-records.

  • khsaa-open-records

    Not yet askedchecked 2026-09-21

    KHSAA holds the records; Arbiter merely hosts them, and a vendor's terms govern the vendor's website rather than the association's records. KHSAA is subject to Kentucky's open government statutes. The catch is standing: KRS 61.872 gives enforceable rights to residents of the Commonwealth, so a non-resident may ask but cannot compel. Worth asking anyway — agencies fill voluntary requests all the time. See docs/get-a-real-season.md.

  • playonsports

    Not yet askedchecked 2026-09-22

    The company that owns MaxPreps, and the reason scraping it was always the wrong door. Their own marketing says state associations keep 'a central data repository to house their member schools' scores and statistics' and that media outlets can 'access reliable, aggregated data around scores, schedules, statistics, and rankings'. That is this project's entire missing dataset, sold rather than fenced. playonsports.com/contact25.

  • sblive

    Refusedchecked 2026-09-24

    scorebooklive.com and sblivesports.com answer CloudFront 403 to this crawler — on /robots.txt itself, before any page. There is no permission to read because the file stating it cannot be reached, and a 403 is an access control. The documented way past it is a residential proxy pool, which is circumvention rather than collection. SBLive powers a good share of the state association scoreboards, so it is the right door; it opens from their side, as a partner.

  • gofan

    Refusedchecked 2026-09-24

    gofan.co/robots.txt: 'User-agent: * / Disallow: /'. Googlebot, Bingbot and Twitterbot are allowed by name and nothing else is. They also disallow a literal Chrome user-agent string, which is a site that has already met the workaround and written it into the file. The blanket disallow is the answer; the Chrome line is them saying they meant it.

  • hudl

    Refusedchecked 2026-09-24

    The clearest refusal found yet, and the sharpest illustration of why robots cannot be the test: hudl.com/robots.txt is 'Allow: /' with four administrative disallows. The terms then prohibit "any automated means, including bots, scrapers, crawlers, or artificial intelligence tools, to access, collect, or extract data from the Services" — and separately forbid using their data to train a model. Permissive robots, total contractual refusal.

  • calhisports

    May collectchecked 2026-09-24

    The first game-level source found whose robots.txt permits this crawler and whose terms do not appear to forbid the use: 'Disallow: /wp-admin/' and 'Crawl-delay: 10', nothing else. Their policy asserts copyright over the content 'and the selection and arrangement thereof', which is the right claim and does not reach a score: who won on Friday is a fact, and facts are not copyrightable however much work went into collecting them. The arrangement is theirs and is not taken: the collector reads one sentence per entry of the free weekly scoreboard post, the result, and nothing of the writing, the rankings or the commentary. Promoted from CANDIDATE to OPEN on 2026-09-24 by Eddie McNichols, on that reading — a decision with a name on it, as the enum asks. The courtesy note (docs/calhisports_email_ready_to_send.txt) says exactly this and promises credit on every score and to stop the moment they ask: credit is the source_url on every row, and stopping is this line set back to FORBIDDEN. Crawl-delay: 10 is honoured by the client, per host, robots fetch included. Gold Club posts are paywalled and are treated as a gate, not parsed.

  • vnn

    Refusedchecked 2026-09-24

    Varsity News Network is now PlayOn — vnnsports.net serves a HubSpot page reading 'VNN is now part of PlayOn'. It is not a separate door; it is the MaxPreps door with a different sign. See playonsports, which is where that conversation happens.

  • social-posts

    Refusedchecked 2026-09-24

    Scraping X/Twitter for coaches posting quarter scores was suggested as the fastest live pipe, and it probably is. It is also forbidden by their terms without a licence, the tooling named for it exists to evade that, and the accounts posting are often students. The lawful version of this idea is somebody at the game typing the score in themselves — which is the 'contributions' source at the top of this file, and is the one live pipe here that needs nobody's permission.

  • maxpreps

    Refusedchecked 2026-09-21

    Terms of Use: "Use any robot, spider, scraper ... to access, monitor, or copy the Platform or its Content or data", and separately forbid "disseminating, distributing, displaying ... any part of the Platform or Content". Their robots.txt now PERMITS football paths, which is why robots cannot be the test.

  • maxpreps-playonsports-marketing

    Refusedchecked 2026-09-22

    maxpreps.playonsports.com is PlayOn Sports' corporate marketing site, HubSpot-hosted, selling advertising and streaming. Its robots.txt permits almost everything because there is nothing on it to protect: no scores, no schedules, no scores to cross-reference. Listed so that a permissive robots.txt on a MaxPreps domain is not mistaken for a way in. The data is at playonsports, above, and it is bought.

  • dave-campbells

    Not yet askedchecked 2026-09-22

    Terms: subscribers "will not copy, publish, or in any way make available publicly any news ... or any other information" without written permission. Scores are that information — but the clause names its own unlock, and nobody has asked. Filed FORBIDDEN until 2026-09-22, which was the error this enum's own docstring warns about: nothing here is refused, it is unrequested. robots.txt disallows only /admin/, /bin/ and other plumbing, so the block is contractual rather than technical, and a contract can be renegotiated. Request drafted at ~/Desktop/DaveCampbells_permission_request.txt; unsent, because the Gmail integration here is read-only. They are the authority for Texas high school football and the single best source in the country for this site's first state.

  • uil

    Refusedchecked 2026-09-21

    Nothing to take. UIL publishes district alignments and a week calendar; /football/scoreboard is a 404 and no page carries a game-level result.

  • school-district-sites

    Refusedchecked 2026-09-22

    Checked 2026-09-22 across Allen, Prosper, Katy and Duncanville ISD. Public bodies, and robots.txt reads permissively — 'User-agent: * / Allow: /'. The trap is one line below it: 'Disallow: /api/'. The athletics pages are JavaScript shells (Allen's is 305 characters of nothing) and every result sits behind the /api/ path that same file refuses. So the permission covers an empty document and the refusal covers the data. Katy carries a second 'User-agent: *' group as well — the MSHSAA duplicate-group trap, which urllib.robotparser does not merge. Listed so nobody spends a week on every district site in the roster.

  • mshsaa

    Refusedchecked 2026-09-21

    The only association in the country that renders results as ordinary HTML, and its robots.txt closes with a second 'User-agent: *' group carrying 'disallow: /'. See scrapers.mshsaa.

  • star-telegram.com

    Refusedchecked 2026-09-24

    Chain terms (mcclatchy.com/terms-of-service, effective 2026-04-14): '10. Use or launch any automated means including spiders, robots, crawlers, scrapers and the like, to download or copy data or content from the Platforms.'; '4.2 ... solely for your personal, non-commercial use ... you will not store or archive a significant portion of the Content or create a database using the Content'; '4.3 Reuse & Republication of Content: You will not reuse, republish or otherwise distribute the Content ... without the express written permission of McClatchy'. The named unlock, mcclatchyreprints.com, is a paid reprint desk rather than a data arrangement, so this is FORBIDDEN rather than LICENSABLE. Independently, every McClatchy host resets the connection for the honest user agent before a byte is served — 'curl: (92) HTTP/2 stream 1 was not closed cleanly: INTERNAL_ERROR' twice, then a 40-45 second HTTP/1.1 timeout with 0 bytes — so robots.txt itself could not be read. Twice refused. This paper: no robots.txt could be retrieved (HTTP/2 INTERNAL_ERROR twice, HTTP/1.1 timed out after 40s with 0 bytes) and /sports/high-school/ failed the same way; no roundup URL obtained.

  • houstonchronicle.com

    Refusedchecked 2026-09-24

    robots.txt permits the high-school paths ('User-agent: * / Disallow: /api / Disallow: /adtest / Disallow: /search / Disallow: /413gkwMT/ ...' and a group of named AI agents including claudebot, claude-web and claude-user with Disallow: /), but the site never lets the honest user agent reach a page: the terms of use, the corporate copy at hearst.com (HTTP 403) and the high-school section all return a 3,038-byte JavaScript 'Client Challenge' interstitial. A challenge is an access control aimed at exactly this kind of client; passing it would take a headless browser or a disguised user agent, which is circumvention. The terms were never read, so the contract question is unanswered: not assessed to a yes. FORBIDDEN for every Hearst paper, dallasnews.com included since its 2025 acquisition (identical robots template, identical challenge). This paper: /terms_of_use/ and /texas-sports-nation/high-school/ both returned the 3,038-byte 'Client Challenge' page; no roundup URL obtained.

  • dallasnews.com

    Refusedchecked 2026-09-24

    robots.txt permits the high-school paths ('User-agent: * / Disallow: /api / Disallow: /adtest / Disallow: /search / Disallow: /413gkwMT/ ...' and a group of named AI agents including claudebot, claude-web and claude-user with Disallow: /), but the site never lets the honest user agent reach a page: the terms of use, the corporate copy at hearst.com (HTTP 403) and the high-school section all return a 3,038-byte JavaScript 'Client Challenge' interstitial. A challenge is an access control aimed at exactly this kind of client; passing it would take a headless browser or a disguised user agent, which is circumvention. The terms were never read, so the contract question is unanswered: not assessed to a yes. FORBIDDEN for every Hearst paper, dallasnews.com included since its 2025 acquisition (identical robots template, identical challenge). This paper: DallasNews Corp was acquired by Hearst in 2025 and serves the same robots template and the same 'Client Challenge' page; /terms-of-service/ and /high-school-sports/football/ both answered with it.

  • wacotrib.com

    Refusedchecked 2026-09-24

    Chain terms (/terms/ on every Lee paper; read from tulsaworld.com and journalnow.com): you agree not to 'Access or attempt to access the Site except as expressly permitted in these Terms of Use; ... Use automated scripts to collect information from, or otherwise interact with, the Site', and 'This Site and Our Content may not be copied, reproduced, republished, uploaded, posted, sold, leased, licensed, sublicensed, transmitted, or distributed without our written permission, except that you may download, display, and print one copy of Our Content on a single computer for your personal, non-commercial use only.' robots.txt on the TownNews/BLOX template disallows only CMS plumbing, so the crawl is technically permitted and contractually forbidden — the MaxPreps shape. No permissions route is named, and the platform redirects the honest user agent's second request to a 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). FORBIDDEN for every Lee paper. This paper: the one Texas host that rendered a page. A weekly 'Central Texas high school football scoreboard' sits at /sports/high-school/football/article_<id>.html and robots.txt permits it (BlueConic meter, section rendered in full); the Lee terms quoted above forbid collecting and republishing it.

  • ajc.com

    Refusedchecked 2026-09-24

    The most explicitly drafted refusal in the survey. AJC Online Services Terms of Use address 'digital engines of any kind, including, without limitation, ones that crawl, index, scrape, copy, store, or transmit digital content' by name, then have every visitor agree it '(i) Will not monitor, gather, copy, or distribute the Content (except as may be a result of standard search engine activity or use of a standard browser) on the Services by using any robot, rover, "bot," spider, scraper, crawler, spyware, engine, device, software, extraction tool, or any other automatic device, utility, or manual process of any kind', bar 'any political or commercial purpose', and bar 'any text or data mining activities'. robots.txt disallows only /search, /google-search and /flatpage-, so the crawl of /sports/varsity/football/scores/ — the best Georgia scoreboard — would be technically permitted and contractually forbidden. Metered (Piano). No permissions route named.

  • savannahnow.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: served the section page to the honest user agent; weekly Friday scoreboard at /story/sports/high-school/football/YYYY/MM/DD/<slug>/<id>/ (e.g. 2026/09/18 'savannah-georgia-high-school-football-scores-live-updates-for-week-5-ghsa'), listed in /news-sitemap.xml. Metered/premium.

  • macon.com

    Refusedchecked 2026-09-24

    Chain terms (mcclatchy.com/terms-of-service, effective 2026-04-14): '10. Use or launch any automated means including spiders, robots, crawlers, scrapers and the like, to download or copy data or content from the Platforms.'; '4.2 ... solely for your personal, non-commercial use ... you will not store or archive a significant portion of the Content or create a database using the Content'; '4.3 Reuse & Republication of Content: You will not reuse, republish or otherwise distribute the Content ... without the express written permission of McClatchy'. The named unlock, mcclatchyreprints.com, is a paid reprint desk rather than a data arrangement, so this is FORBIDDEN rather than LICENSABLE. Independently, every McClatchy host resets the connection for the honest user agent before a byte is served — 'curl: (92) HTTP/2 stream 1 was not closed cleanly: INTERNAL_ERROR' twice, then a 40-45 second HTTP/1.1 timeout with 0 bytes — so robots.txt itself could not be read. Twice refused. This paper: robots.txt and /sports/high-school/ both failed with HTTP/2 INTERNAL_ERROR and 0 bytes; no roundup URL obtained.

  • gwinnettdailypost.com

    Refusedchecked 2026-09-24

    robots.txt permits /sports/ (only TownNews plumbing is disallowed; named AI agents get Disallow: /), but the 'Terms of Use' link (/site/terms.html) redirected: the host redirected the honest user agent to a TownNews 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign) on the second request. The terms were never read, so the reuse question is unanswered — not assessed to a yes — and a captcha is a 'no' to automated clients. Football section /sports/prep/sport/football/; metered (TownNews access-offers modal).

  • cleveland.com

    Refusedchecked 2026-09-24

    Chain terms (advancelocal.com User Agreement, updated 2024-08-01): you may not 'use any bots, cheats, macros, scripts ... or use any other automated process, or engage in meta-searching or periodic caching of information, to access, visit and/or use the Service', nor 'copy, harvest, crawl, index, scrape, spider, mine, gather, extract, compile, obtain, aggregate, capture, access, store, or republish any Content on or through the Service, including by an automated or manual process or otherwise, for any and all purposes other than indexing Content for inclusion in a Search Engine', and 'You' is defined to include 'digital engines of any kind that harvest, crawl, index, scrape, spider, or mine digital content'. A flat prohibition, not one conditioned on permission. Independently, every Advance host answers HTTP 403 through DataDome ('Please enable JS and disable any ad blocker', geo.captcha-delivery.com) to the honest user agent on the section page and on the user agreement itself. robots.txt permits /highschoolsports/, which is necessary and not sufficient. FORBIDDEN for every Advance Local site. This paper: /highschoolsports/ and /user-agreement/ both returned HTTP 403 from DataDome; no roundup URL obtained.

  • dispatch.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: served the section page; the fullest Ohio weekly scoreboard of anything surveyed, at /story/sports/high-school/football/2026/09/19/ohio-high-school-football-ohsaa-week-5-games-scores-stats/91800433007/ and a Friday live version, listed in /news-sitemap.xml. Metered/premium (Gannett meter plus Sophi).

  • daytondailynews.com

    Refusedchecked 2026-09-24

    robots.txt permits /sports/ and the server never lets a plain client see one: /terms/ and /sports/high-school/ both redirected, and the host redirected the honest user agent to a TownNews 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). The terms behind it were never read, so the contract question stands unanswered — not assessed to a yes. Cox's Ohio papers (Journal-News, Springfield News-Sun) share the stack; Cox's AJC has its own explicit refusal.

  • toledoblade.com

    Refusedchecked 2026-09-24

    The friendliest robots.txt in the survey ('# Be nice.', a dozen legacy Our-Town disallows, nothing touching /sports/) on a site whose pages are empty without JavaScript: /terms-of-service and /terms both returned 55-59 KB whose only text is the navigation, and /sports/high-school carries no article links in the served HTML. The terms could not be read, so whether reuse is permitted is unknown — not assessed to a yes. Metered (Sophi, Piano). Worth a human reading the terms in a browser: if permissive, the one Ohio outlet that could move to CANDIDATE.

  • tampabay.com

    Refusedchecked 2026-09-24

    robots.txt is the most permissive of the Florida set for a generic agent ('User-agent: * / Allow: /ads.txt' and nothing else) while blocking eight named AI crawlers outright. The terms of use could not be found: /terms/ and /terms-of-use/ both 404 and the footer that would link them is rendered by JavaScript. No clause to read means the contract question is unanswered, and this register does not promote an outlet on the absence of a clause — not assessed to a yes. Section /sports/high-schools/ (Arc; stories hydrated client-side); metered (Sophi + Zephr).

  • miamiherald.com

    Refusedchecked 2026-09-24

    Chain terms (mcclatchy.com/terms-of-service, effective 2026-04-14): '10. Use or launch any automated means including spiders, robots, crawlers, scrapers and the like, to download or copy data or content from the Platforms.'; '4.2 ... solely for your personal, non-commercial use ... you will not store or archive a significant portion of the Content or create a database using the Content'; '4.3 Reuse & Republication of Content: You will not reuse, republish or otherwise distribute the Content ... without the express written permission of McClatchy'. The named unlock, mcclatchyreprints.com, is a paid reprint desk rather than a data arrangement, so this is FORBIDDEN rather than LICENSABLE. Independently, every McClatchy host resets the connection for the honest user agent before a byte is served — 'curl: (92) HTTP/2 stream 1 was not closed cleanly: INTERNAL_ERROR' twice, then a 40-45 second HTTP/1.1 timeout with 0 bytes — so robots.txt itself could not be read. Twice refused. This paper: robots.txt and /sports/high-school/ both failed with HTTP/2 INTERNAL_ERROR and 0 bytes; no roundup URL obtained.

  • orlandosentinel.com

    Refusedchecked 2026-09-24

    Chain terms (tribpub.com/central-terms-of-service/, updated 2024-11-08, one document for the Orlando Sentinel, Chicago Tribune, Sun Sentinel, Baltimore Sun, Hartford Courant, Daily Press, Virginian-Pilot, Morning Call and New York Daily News): 'You may not scrape or otherwise copy our Content without our permission'; you agree not to 'use robots, spiders, scripts, service, software, or any manual or automatic device, tool, or process designed to data mine or scrape the Content' nor to access 'the Site using automated means (such as harvesting bots, robots, spiders, or scrapers) without our prior permission', and 'you may not republish any portion of the Content ... or incorporate the Content in any database, compilation, archive, cache, or similar medium without Tribune Publishing's prior written consent.' The Tribune's robots.txt header adds 'Any other uses are prohibited, including ... (3) caching or archiving the Content; and/or (4) any commercial purposes.' robots.txt for * blocks only WordPress plumbing; the contract forbids the crawl three ways. Both assessors filed it FORBIDDEN and so does this entry, with the note that the terms name a route — termsofservice@tribpub.com — so a letter is possible. This paper: served the section page; a weekly Friday live FHSAA scoreboard at /2026/09/24/orlando-area-week-6-live-football-scoreboard-fhsaa-district-races-heat-up/ (WordPress /YYYY/MM/DD/<slug>/). The working section is /sports/high-school-sports/ (/sports/high-school/ redirects to a 2001 archive article). Metered with Auth0 regwall.

  • jacksonville.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: served the section page; Friday scores roundup at /story/sports/high-school/football/2026/09/18/northeast-florida-high-school-football-week-5-scores/87976478007/ plus a statewide FHSAA version and a live board, listed in /news-sitemap.xml. Metered/premium.

  • latimes.com

    Not yet askedchecked 2026-09-24

    robots.txt permits /sports/highschool (disallows only /search, /config/, /test/ and the like; no Crawl-delay). Terms of Service: 'You may not scrape or otherwise copy our Content without our permission.'; 'You agree not to ... use any data mining, data gathering or extraction method.'; 'you may not republish any portion of the Content on any Internet, Intranet, extranet site ... or incorporate the Content in any database, compilation, archive, cache, or similar medium'; 'solely for your personal, non-commercial use'. The contract forbids automated access today. LICENSABLE rather than FORBIDDEN because the clause is conditioned on 'our permission' and the site served every page to the honest user agent: a written ask could open it. Roundups at /sports/highschool/story/YYYY-MM-DD/high-school-football-{thursdays,fridays,saturdays}-scores, in /news-sitemap.xml. Metered. Nothing may be fetched until granted.

  • sfchronicle.com

    Refusedchecked 2026-09-24

    robots.txt permits the high-school paths ('User-agent: * / Disallow: /api / Disallow: /adtest / Disallow: /search / Disallow: /413gkwMT/ ...' and a group of named AI agents including claudebot, claude-web and claude-user with Disallow: /), but the site never lets the honest user agent reach a page: the terms of use, the corporate copy at hearst.com (HTTP 403) and the high-school section all return a 3,038-byte JavaScript 'Client Challenge' interstitial. A challenge is an access control aimed at exactly this kind of client; passing it would take a headless browser or a disguised user agent, which is circumvention. The terms were never read, so the contract question is unanswered: not assessed to a yes. FORBIDDEN for every Hearst paper, dallasnews.com included since its 2025 acquisition (identical robots template, identical challenge). This paper: /terms_of_use/ and /sports/highschool/ both returned the 'Client Challenge' page; hearst.com terms returned 403; no roundup URL obtained.

  • sacbee.com

    Refusedchecked 2026-09-24

    Chain terms (mcclatchy.com/terms-of-service, effective 2026-04-14): '10. Use or launch any automated means including spiders, robots, crawlers, scrapers and the like, to download or copy data or content from the Platforms.'; '4.2 ... solely for your personal, non-commercial use ... you will not store or archive a significant portion of the Content or create a database using the Content'; '4.3 Reuse & Republication of Content: You will not reuse, republish or otherwise distribute the Content ... without the express written permission of McClatchy'. The named unlock, mcclatchyreprints.com, is a paid reprint desk rather than a data arrangement, so this is FORBIDDEN rather than LICENSABLE. Independently, every McClatchy host resets the connection for the honest user agent before a byte is served — 'curl: (92) HTTP/2 stream 1 was not closed cleanly: INTERNAL_ERROR' twice, then a 40-45 second HTTP/1.1 timeout with 0 bytes — so robots.txt itself could not be read. Twice refused. This paper: robots.txt was served (176 lines, nothing under /sports; Crawl-delay 2 and 3 appear) but /sports/high-school/ hung to a 40-second read timeout on two attempts, which is a soft refusal and is treated as one.

  • mercurynews.com

    Not yet askedchecked 2026-09-24

    Chain terms (medianewsgroup.com/terms-of-use/): '(iii) use any robots, spiders, crawlers, data mining or extraction technology, or other automated scripts or means to collect or scrape information from or otherwise interact with the Technology or copy our Content without written permission;' and 'Any other use of Content, including without limitation republication of any Content or any commercial or public use, requires the prior written consent of MediaNews Group.' robots.txt permits the section and declares Crawl-delay: 10, which the reader would honour. LICENSABLE rather than FORBIDDEN because both clauses are conditioned on written permission and the hosts served every page to the honest user agent: a written ask to MediaNews Group is the door, and nothing may be fetched until it is granted. One entry covers every MediaNews/Alden paper (Mercury News, East Bay Times, OC Register, Press-Enterprise, San Diego Union-Tribune, Denver Post, Boston Herald, Trentonian, Reading Eagle and others). This paper: served the section page; weekly 'Bay Area prep football week N weekend scoreboard' at /2026/09/20/bay-area-prep-football-week-4-2026-weekend-scoreboard-how-top-25-fared/ and a Friday roundup, from /sports/high-school-sports/ and /sitemap.xml. Metered with registration wall. Crawl-delay: 10 declared.

  • nola.com

    Refusedchecked 2026-09-24

    robots.txt does not disallow /sports/high_school/ (161 lines of TownNews plumbing; the only sports line is /sports/scott_rabalais/), but the server answered the honest user agent's very first content requests — /terms/ and the section page, ten seconds apart — with HTTP 429 'Too Many Requests client_ip: <ip> request_id: <id>'. A rate limit on request one is a refusal of this crawler, and the terms could not be read — not assessed to a yes. Sister paper theadvocate.com additionally disallows its prep-scores paths.

  • theadvocate.com

    Refusedchecked 2026-09-24

    robots.txt explicitly disallows the scoreboard pages: 'Disallow: /acadiana/sports/high_schools/prep_scores/' and 'Disallow: /prepscores/acadiana/'. That alone decides it — the one thing this reader would read is the thing they excluded. On top of that /terms/ answered the first request with HTTP 429 'Too Many Requests'. Forbidden on robots, unreadable on terms.

  • shreveporttimes.com

    Refusedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: the host refuses the honest user agent outright — /robots.txt itself returned HTTP 403 '<title>Access Restricted</title> ... We're unable to allow access to this page at this time ... {RM:000}', as did /sports/high-school/. A 403 is an answer and the answer is stop, whatever the chain's permissions desk would say; the chain terms page (cm.shreveporttimes.com/terms/) is a 203-character JavaScript shell to this client. FORBIDDEN on its own account.

  • americanpress.com

    Refusedchecked 2026-09-24

    robots.txt allows everything ('User-agent: * / Disallow:'), and the Terms of Service (/services/terms-of-service/) say 'You may not modify, copy, reproduce, republish, upload, post, transmit or distribute in any way any material from the Service', 'for your personal, non-commercial use only' and 'You agree to use the Service ... for legitimate, non-commercial purposes only.' No robot clause, and a final score is a fact rather than 'material', but Hometown Line is a commercial product and its use is outside the licence on its face; the republication clause is flat rather than conditioned on permission. When in doubt, forbidden. Weekly SWLA prep coverage at /YYYY/MM/DD/<slug>/, in /sitemap-news.xml; metered. A written ask to Carpenter Media Group could reopen it.

  • nj.com

    Refusedchecked 2026-09-24

    Chain terms (advancelocal.com User Agreement, updated 2024-08-01): you may not 'use any bots, cheats, macros, scripts ... or use any other automated process, or engage in meta-searching or periodic caching of information, to access, visit and/or use the Service', nor 'copy, harvest, crawl, index, scrape, spider, mine, gather, extract, compile, obtain, aggregate, capture, access, store, or republish any Content on or through the Service, including by an automated or manual process or otherwise, for any and all purposes other than indexing Content for inclusion in a Search Engine', and 'You' is defined to include 'digital engines of any kind that harvest, crawl, index, scrape, spider, or mine digital content'. A flat prohibition, not one conditioned on permission. Independently, every Advance host answers HTTP 403 through DataDome ('Please enable JS and disable any ad blocker', geo.captcha-delivery.com) to the honest user agent on the section page and on the user agreement itself. robots.txt permits /highschoolsports/, which is necessary and not sufficient. FORBIDDEN for every Advance Local site. This paper: /highschoolsports/football/ returned HTTP 403 to the honest user agent; no roundup URL obtained.

  • northjersey.com

    Refusedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: /robots.txt and /sports/high-school/football/ both returned HTTP 403 'Access Restricted' to the honest user agent, and cm.northjersey.com/terms/ is a 218-character client-rendered shell. A 403 is an answer: FORBIDDEN on its own account, whatever the chain's permissions desk would say.

  • pressofatlanticcity.com

    Refusedchecked 2026-09-24

    Chain terms (/terms/ on every Lee paper; read from tulsaworld.com and journalnow.com): you agree not to 'Access or attempt to access the Site except as expressly permitted in these Terms of Use; ... Use automated scripts to collect information from, or otherwise interact with, the Site', and 'This Site and Our Content may not be copied, reproduced, republished, uploaded, posted, sold, leased, licensed, sublicensed, transmitted, or distributed without our written permission, except that you may download, display, and print one copy of Our Content on a single computer for your personal, non-commercial use only.' robots.txt on the TownNews/BLOX template disallows only CMS plumbing, so the crawl is technically permitted and contractually forbidden — the MaxPreps shape. No permissions route is named, and the platform redirects the honest user agent's second request to a 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). FORBIDDEN for every Lee paper. This paper: /terms/, /sports/high-school/ and the corporate lee.net/terms-of-use/ all answered the first request with HTTP 429 'Too Many Requests client_ip: <ip> request_id: <id>'.

  • trentonian.com

    Not yet askedchecked 2026-09-24

    Chain terms (medianewsgroup.com/terms-of-use/): '(iii) use any robots, spiders, crawlers, data mining or extraction technology, or other automated scripts or means to collect or scrape information from or otherwise interact with the Technology or copy our Content without written permission;' and 'Any other use of Content, including without limitation republication of any Content or any commercial or public use, requires the prior written consent of MediaNews Group.' robots.txt permits the section and declares Crawl-delay: 10, which the reader would honour. LICENSABLE rather than FORBIDDEN because both clauses are conditioned on written permission and the hosts served every page to the honest user agent: a written ask to MediaNews Group is the door, and nothing may be fetched until it is granted. One entry covers every MediaNews/Alden paper (Mercury News, East Bay Times, OC Register, Press-Enterprise, San Diego Union-Tribune, Denver Post, Boston Herald, Trentonian, Reading Eagle and others). This paper: served the section page (/sports/high-school-sports/, footer links the chain terms); publishes per-game stories and a weekly picks column rather than one scoreboard, so a weaker target than the Mercury News. Metered with registration wall. Crawl-delay: 10 declared.

  • inquirer.com

    Not yet askedchecked 2026-09-24

    robots.txt permits /high-school-sports/ (disallows /light/, /search, /sports/betting/, /wires/ and test paths). Terms and Conditions: 'You are also specifically prohibited from gathering, collecting, transmitting and/or using any materials, data or information from any of our Products for any purpose whatsoever, except as otherwise provided in these Terms of Use, without our express written permission, provided in advance. This prohibition applies to, among other things, scraping any of our Products, using cookies and/or any device, tool or software that "spiders" or "crawls"'; and '8. Permitted Uses: ... Any commercial use of any Product ... is prohibited, except with the prior written consent of The Inquirer.' Articles also sit behind 'Keep reading by creating a free account', which is a login: stop. LICENSABLE because both clauses are expressly conditioned on advance written permission and the site served pages to the honest user agent. Friday scoreboards at /high-school-sports/<slug>-YYYYMMDD.html via the 48-hour news sitemap. Nothing may be fetched until granted.

  • post-gazette.com

    Refusedchecked 2026-09-24

    robots.txt is permissive ('Disallow: /admin / Disallow: /*.print / Disallow: /zillowarticles'), but /terms, /terms-of-use and /sports/hsfootball all return the same 2,609-character navigation shell — empty title, 'noindex, nofollow', no clause text, no article links — and the section carries Piano/tinypass scripts with 'subscribe to continue'. The terms cannot be read and the content sits behind a hard paywall; either alone is a stop. Not assessed to a yes.

  • pennlive.com

    Refusedchecked 2026-09-24

    Chain terms (advancelocal.com User Agreement, updated 2024-08-01): you may not 'use any bots, cheats, macros, scripts ... or use any other automated process, or engage in meta-searching or periodic caching of information, to access, visit and/or use the Service', nor 'copy, harvest, crawl, index, scrape, spider, mine, gather, extract, compile, obtain, aggregate, capture, access, store, or republish any Content on or through the Service, including by an automated or manual process or otherwise, for any and all purposes other than indexing Content for inclusion in a Search Engine', and 'You' is defined to include 'digital engines of any kind that harvest, crawl, index, scrape, spider, or mine digital content'. A flat prohibition, not one conditioned on permission. Independently, every Advance host answers HTTP 403 through DataDome ('Please enable JS and disable any ad blocker', geo.captcha-delivery.com) to the honest user agent on the section page and on the user agreement itself. robots.txt permits /highschoolsports/, which is necessary and not sufficient. FORBIDDEN for every Advance Local site. This paper: /highschoolsports/football/ returned HTTP 403 to the honest user agent; no roundup URL obtained.

  • triblive.com

    Refusedchecked 2026-09-24

    robots.txt is a bare 'User-agent: *' with no rules — everything permitted — and TribLive runs a purpose-built HSSN Scores page (/sports/hssn/hssn-scores/), which makes it the most natural licensing conversation in Pennsylvania. But the Terms of Service (/termsofservice/) say you agree not to 'Use any robot, spider or other automatic device, process or means to access the Website for any purpose, including monitoring or copying any of the material on the Website', that 'You must not access or use for any commercial purposes any part of the Website', and that 'unless expressly authorized by the Company, in writing, you must not publish, reproduce, distribute, republish, enter into a database ... any part of the content'. The robot clause is unconditional, so this is 'they said no' rather than 'nobody asked': FORBIDDEN. The republication clause does allow written authorisation, so a letter about HSSN scores is worth sending.

  • al.com

    Refusedchecked 2026-09-24

    Chain terms (advancelocal.com User Agreement, updated 2024-08-01): you may not 'use any bots, cheats, macros, scripts ... or use any other automated process, or engage in meta-searching or periodic caching of information, to access, visit and/or use the Service', nor 'copy, harvest, crawl, index, scrape, spider, mine, gather, extract, compile, obtain, aggregate, capture, access, store, or republish any Content on or through the Service, including by an automated or manual process or otherwise, for any and all purposes other than indexing Content for inclusion in a Search Engine', and 'You' is defined to include 'digital engines of any kind that harvest, crawl, index, scrape, spider, or mine digital content'. A flat prohibition, not one conditioned on permission. Independently, every Advance host answers HTTP 403 through DataDome ('Please enable JS and disable any ad blocker', geo.captcha-delivery.com) to the honest user agent on the section page and on the user agreement itself. robots.txt permits /highschoolsports/, which is necessary and not sufficient. FORBIDDEN for every Advance Local site. This paper: /user-agreement/ returned HTTP 403 with a DataDome challenge page, so no article was requested. Weekly statewide AHSAA scoreboards live under /highschoolsports/ (from general knowledge, not verified by fetch).

  • montgomeryadvertiser.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: served the AHSAA scoreboard page (HTTP 200, tagged 'free - free always', isAccessibleForFree:true) at /story/sports/high-school/football/2026/09/24/see-all-ahsaa-week-5-alabama-high-school-football-scores-here/91826228007/ — but the scores are not in the HTML: the page embeds a ScoreStream widget (scorestream.com/widgets/scoreboards/vert?userWidgetId=56437), so the real feed is ScoreStream, already LICENSABLE in this register, not Gannett's markup.

  • dothaneagle.com

    Refusedchecked 2026-09-24

    Chain terms (/terms/ on every Lee paper; read from tulsaworld.com and journalnow.com): you agree not to 'Access or attempt to access the Site except as expressly permitted in these Terms of Use; ... Use automated scripts to collect information from, or otherwise interact with, the Site', and 'This Site and Our Content may not be copied, reproduced, republished, uploaded, posted, sold, leased, licensed, sublicensed, transmitted, or distributed without our written permission, except that you may download, display, and print one copy of Our Content on a single computer for your personal, non-commercial use only.' robots.txt on the TownNews/BLOX template disallows only CMS plumbing, so the crawl is technically permitted and contractually forbidden — the MaxPreps shape. No permissions route is named, and the platform redirects the honest user agent's second request to a 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). FORBIDDEN for every Lee paper. This paper: the 'Friday night football scores from around the Wiregrass and state' roundup request (/sports/high-school/football/article_<id>.html) was redirected: the host redirected the honest user agent to a TownNews 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). No content page was served.

  • timesdaily.com

    Refusedchecked 2026-09-24

    robots.txt permits /sports/ (63 lines of TownNews plumbing), but the only content request — 'Week 3: Alabama high school football scores' at /sports/high_school/week-3-alabama-high-school-football-scores/article_<id>.html — was redirected: the host redirected the honest user agent to a TownNews 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). The terms (/site/terms.html) were therefore not requested; a search snippet, unverified, limits copying to 'personal use only' with other copying 'prohibited without prior written permission'. Unread terms are doubt — not assessed to a yes — and a person must read them in a browser before anyone promotes this.

  • tennessean.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: robots.txt and /news-sitemap.xml were served; the Friday scores post lands under /story/sports/high-school/football/YYYY/MM/DD/ and is discoverable from the sitemap and /sports/high-school/football/. No article fetched.

  • timesfreepress.com

    Not yet askedchecked 2026-09-24

    robots.txt: 'User-agent: * / Allow: /' with a handful of disallows (/assets/, /cgi-bin/, /puzzles/) and 'Content-Signal: ai-train=no, search=yes, ai-input=no'. Terms and Conditions (updated 2026-03-11): 'not to use any data mining, robots, malware, or any data gathering or extraction method in connection with your use of our Digital Products'; '(c) to use any computer program, bot, robot, spider, offline reader, website search/retrieval application, or other manual automatic device, tool, or process to retrieve, index, data mine, or in any way reproduce ... the Content'; 'You may not ... incorporate the Content in any database, compilation, archive or cache.' LICENSABLE because the terms name the door twice: 'Requests to use the Content for any purpose other than as permitted in this Copyright section should be directed to our Permissions' and 'REPRINT PERMISSION: To request permission to republish information, fax or email Alison Gerber, Managing Editor, at 423-756-6900 or agerber@timesfreepress.com.' The 'State prep football scores for Georgia, Tennessee' posts (/news/YYYY/mon/DD/<slug>/) are the best TN/GA score list found; served in full to the honest user agent on first visit, metered (Zephr). Needs a letter, not a crawler.

  • dailymemphian.com

    Not yet askedchecked 2026-09-24

    robots.txt permits (CCBot blocked; /feed/tag/, /start, /ios-beta disallowed for *). The terms (/tos) contain no robot or scraping clause, but: 'You may not obtain or attempt to obtain any materials or information through any means not intentionally made available or provided for through the Site.'; 'You will use protected content solely for your personal use, and will make no other use of the content without the express written permission of Daily Memphian and the copyright owner.'; 'Our site is subscription-based'. Publishing scores credited 'via The Daily Memphian' is a non-personal use, so: permitted only with written permission, which nobody has asked for. There is also no scoreboard list — the Friday 'prep report' is prose, which this project does not copy — so the yield would be a handful of results a week. Metered (Piano).

  • timesnews.net

    Refusedchecked 2026-09-24

    robots.txt permits (disallows /custom/, /images/, /oauth/, /user/login/ and the like), but the terms of use (/live-content/terms-of-use/) are a client-rendered mynews360.com shell with no clause text in the served HTML, and /site/terms.html is a 404. Unread terms are doubt, and doubt is forbidden — not assessed to a yes. Re-evaluate only after a person reads the terms in a browser and quotes them here.

  • charlotteobserver.com

    Refusedchecked 2026-09-24

    Chain terms (mcclatchy.com/terms-of-service, effective 2026-04-14): '10. Use or launch any automated means including spiders, robots, crawlers, scrapers and the like, to download or copy data or content from the Platforms.'; '4.2 ... solely for your personal, non-commercial use ... you will not store or archive a significant portion of the Content or create a database using the Content'; '4.3 Reuse & Republication of Content: You will not reuse, republish or otherwise distribute the Content ... without the express written permission of McClatchy'. The named unlock, mcclatchyreprints.com, is a paid reprint desk rather than a data arrangement, so this is FORBIDDEN rather than LICENSABLE. Independently, every McClatchy host resets the connection for the honest user agent before a byte is served — 'curl: (92) HTTP/2 stream 1 was not closed cleanly: INTERNAL_ERROR' twice, then a 40-45 second HTTP/1.1 timeout with 0 bytes — so robots.txt itself could not be read. Twice refused. This paper: /robots.txt never completed for the honest user agent (HTTP/2 INTERNAL_ERROR twice, 45-second HTTP/1.1 timeout with 0 bytes); nothing further was requested.

  • highschoolot.com

    Refusedchecked 2026-09-24

    robots.txt: 'User-agent: * / Disallow: /' (lines 91-92; two dozen named search and social agents get their own allow groups above it). A complete refusal to any crawler not named, and a robots allow is necessary before anything else is considered, so the terms were not even requested. The best NC statewide scoreboard (crowd-reported to HighSchoolOT@wral.com / #HSOTscores) and it is closed to us; the only path is a data arrangement with WRAL/Capitol Broadcasting, which nobody has sought, but with a blanket disallow and unread terms this is forbidden, not licensable.

  • journalnow.com

    Refusedchecked 2026-09-24

    Chain terms (/terms/ on every Lee paper; read from tulsaworld.com and journalnow.com): you agree not to 'Access or attempt to access the Site except as expressly permitted in these Terms of Use; ... Use automated scripts to collect information from, or otherwise interact with, the Site', and 'This Site and Our Content may not be copied, reproduced, republished, uploaded, posted, sold, leased, licensed, sublicensed, transmitted, or distributed without our written permission, except that you may download, display, and print one copy of Our Content on a single computer for your personal, non-commercial use only.' robots.txt on the TownNews/BLOX template disallows only CMS plumbing, so the crawl is technically permitted and contractually forbidden — the MaxPreps shape. No permissions route is named, and the platform redirects the honest user agent's second request to a 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). FORBIDDEN for every Lee paper. This paper: served the terms (this is where the Lee clause was read) and the 'High School Football scores for Week 5' article at /sports/high-school/football/article_<id>.html — whose body is subscriber-only encrypted text ('<div class="subscriber-only encrypted-content lee-article-text" style="display:none">', isAccessibleForFree=false). A paywall is a stop on its own.

  • fayobserver.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: robots.txt and /news-sitemap.xml were served; Friday scores posts follow /story/sports/high-school/YYYY/MM/DD/ and are discoverable from the sitemap and /sports/high-school/. No article fetched.

  • postandcourier.com

    Refusedchecked 2026-09-24

    robots.txt (659 lines, nothing under /sports/) permits, but the one content request — 'Updated high school football scores' at /sports/updated-high-school-football-scores/article_<id>.html — was redirected: the host redirected the honest user agent to a TownNews 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). The posture says stop, so the terms (/site/terms.html) were not requested and remain unread: not assessed to a yes.

  • thestate.com

    Refusedchecked 2026-09-24

    Chain terms (mcclatchy.com/terms-of-service, effective 2026-04-14): '10. Use or launch any automated means including spiders, robots, crawlers, scrapers and the like, to download or copy data or content from the Platforms.'; '4.2 ... solely for your personal, non-commercial use ... you will not store or archive a significant portion of the Content or create a database using the Content'; '4.3 Reuse & Republication of Content: You will not reuse, republish or otherwise distribute the Content ... without the express written permission of McClatchy'. The named unlock, mcclatchyreprints.com, is a paid reprint desk rather than a data arrangement, so this is FORBIDDEN rather than LICENSABLE. Independently, every McClatchy host resets the connection for the honest user agent before a byte is served — 'curl: (92) HTTP/2 stream 1 was not closed cleanly: INTERNAL_ERROR' twice, then a 40-45 second HTTP/1.1 timeout with 0 bytes — so robots.txt itself could not be read. Twice refused. This paper: /robots.txt never completed for the honest user agent (HTTP/2 INTERNAL_ERROR twice, 45-second HTTP/1.1 timeout with 0 bytes); nothing further was requested.

  • greenvilleonline.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: robots.txt and /news-sitemap.xml were served; Friday scores posts follow /story/sports/high-school/YYYY/MM/DD/ and are discoverable from the sitemap and /sports/high-school/. No article fetched.

  • indexjournal.com

    Refusedchecked 2026-09-24

    robots.txt permits /sports/, but the terms request itself (/site/terms.html) was redirected: the host redirected the honest user agent to a TownNews 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). The terms are unread and the host has shown it will gate the honest user agent: not assessed to a yes. Scores page /sports/lakelands/scores/ (from search, not fetched).

  • clarionledger.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: served 'Mississippi high school football scores for MHSAA 2026 Week 5' at /story/sports/high-school/2026/09/24/mississippi-high-school-football-scores-mhsaa-2026-week-5/91861046007/ (HTTP 200, 'free - free always', isAccessibleForFree:true) as plain HTML with no widget — the most useful Gannett page found, and still not open pending a request.

  • sunherald.com

    Refusedchecked 2026-09-24

    Chain terms (mcclatchy.com/terms-of-service, effective 2026-04-14): '10. Use or launch any automated means including spiders, robots, crawlers, scrapers and the like, to download or copy data or content from the Platforms.'; '4.2 ... solely for your personal, non-commercial use ... you will not store or archive a significant portion of the Content or create a database using the Content'; '4.3 Reuse & Republication of Content: You will not reuse, republish or otherwise distribute the Content ... without the express written permission of McClatchy'. The named unlock, mcclatchyreprints.com, is a paid reprint desk rather than a data arrangement, so this is FORBIDDEN rather than LICENSABLE. Independently, every McClatchy host resets the connection for the honest user agent before a byte is served — 'curl: (92) HTTP/2 stream 1 was not closed cleanly: INTERNAL_ERROR' twice, then a 40-45 second HTTP/1.1 timeout with 0 bytes — so robots.txt itself could not be read. Twice refused. This paper: /robots.txt never completed for the honest user agent (HTTP/2 INTERNAL_ERROR twice, 45-second HTTP/1.1 timeout with 0 bytes); nothing further was requested.

  • djournal.com

    Not yet askedchecked 2026-09-24

    robots.txt permits /sports/ (90 lines, newsletters and TownNews plumbing disallowed). Terms (/site/terms.html): 'Use or launch any automated means including spiders, robots, crawlers, scrapers and the like, to download or copy data or content from DJournal.com.' is a prohibited act; 'You may not reproduce, republish or redistribute Content ... without the written consent of the copyright owner'; '5.2 ... solely for your personal, non-commercial use ... you will not store or archive a significant portion of the Content or create a database using the Content'. LICENSABLE because the terms name a specific permissions route — 'please call 662-842-2611, or send an e-mail to editor@journalinc.com' — and the host served the honest user agent without a challenge. Weekly 'Week N Northeast Mississippi high school football scores' posts under /sports/high-school/ are exactly the kind of thing worth asking for.

  • vicksburgpost.com

    Refusedchecked 2026-09-24

    robots.txt says 'Allow: /' (one query-string disallow), and the host then answered the first real request, /terms, with HTTP 403 from CloudFront ('ERROR: The request could not be satisfied ... Request blocked.'). A 403 is an answer; the terms are unread — not assessed to a yes. Even if a person reads the terms and finds no automated-access clause, the CDN block stands as the outlet's operative answer to this client.

  • chicagotribune.com

    Refusedchecked 2026-09-24

    Chain terms (tribpub.com/central-terms-of-service/, updated 2024-11-08, one document for the Orlando Sentinel, Chicago Tribune, Sun Sentinel, Baltimore Sun, Hartford Courant, Daily Press, Virginian-Pilot, Morning Call and New York Daily News): 'You may not scrape or otherwise copy our Content without our permission'; you agree not to 'use robots, spiders, scripts, service, software, or any manual or automatic device, tool, or process designed to data mine or scrape the Content' nor to access 'the Site using automated means (such as harvesting bots, robots, spiders, or scrapers) without our prior permission', and 'you may not republish any portion of the Content ... or incorporate the Content in any database, compilation, archive, cache, or similar medium without Tribune Publishing's prior written consent.' The Tribune's robots.txt header adds 'Any other uses are prohibited, including ... (3) caching or archiving the Content; and/or (4) any commercial purposes.' robots.txt for * blocks only WordPress plumbing; the contract forbids the crawl three ways. Both assessors filed it FORBIDDEN and so does this entry, with the note that the terms name a route — termsofservice@tribpub.com — so a letter is possible. This paper: served the section page (/high-school-sports/, metered, paywallId 'paywall-sspw-mg2pw01-fg24'); publishes game stories and area reports at /YYYY/MM/DD/<slug>/ rather than one statewide list.

  • chicago.suntimes.com

    Refusedchecked 2026-09-24

    robots.txt permits the scores pages (disallows /search, /pages/sponsored, /_track). Terms of Use (/legal/terms-of-use): you may not do anything that 'Collects or attempts to collect any user content or information, or otherwise accesses the Websites using automated means (such as harvesting bots, robots, spiders, or scrapers) without our express, prior written permission'; 'you may not republish any portion of the Content ... or incorporate the Content in any database, compilation, archive, cache, or similar medium'; 'You may not scrape or otherwise copy our Content without our permission.' The weekly statewide list at /high-school-football/scores/YYYY/MM/DD/<slug> is exactly the conference-by-conference final-score list this project wants, served in full behind a Piano meter, and it is one written request away — the clause names the route. Filed FORBIDDEN as assessed, until that request is granted.

  • dailyherald.com

    Refusedchecked 2026-09-24

    robots.txt is the single line 'User-agent: *' with no rules, but the site fronts an Imperva/Incapsula JavaScript challenge that was served to the plain user agent on the very first content request (/scoreboard/: a 212-byte shell with '<META NAME="robots" CONTENT="noindex,nofollow">' and an _Incapsula_Resource script). A challenge page is a gate, and a gate is an answer: stop. The terms were not read from the live site — not assessed to a yes — and an indexed snippet forbids using content 'to construct any kind of database'.

  • shawlocal.com

    Refusedchecked 2026-09-24

    robots.txt permits (disallows /admin/, /examples/, /search/, /test/) and the terms — the 'Terms of Service' section of /privacy/; /terms-of-service/ is a 404 — carry no clause on robots, crawlers or automated access. But: 'Unless otherwise specified, the Service is intended for your personal, noncommercial use only. You may not modify, copy, reproduce, republish, upload, post, transmit or distribute in any way any material, including code and software, from the Service.' Hometown Line is a commercial product that would copy and republish the results, so the licence as written does not cover the use even though a final score is a fact. The strongest ask-first target in Illinois: a weekly statewide list at /friday-night-drive/YYYY/MM/DD/week-N-illinois-high-school-football-scores-for-the-2026-season/, served in full (Piano meter), no bot clause, no gate. Forbidden until Shaw Media says yes in writing.

  • mlive.com

    Refusedchecked 2026-09-24

    Chain terms (advancelocal.com User Agreement, updated 2024-08-01): you may not 'use any bots, cheats, macros, scripts ... or use any other automated process, or engage in meta-searching or periodic caching of information, to access, visit and/or use the Service', nor 'copy, harvest, crawl, index, scrape, spider, mine, gather, extract, compile, obtain, aggregate, capture, access, store, or republish any Content on or through the Service, including by an automated or manual process or otherwise, for any and all purposes other than indexing Content for inclusion in a Search Engine', and 'You' is defined to include 'digital engines of any kind that harvest, crawl, index, scrape, spider, or mine digital content'. A flat prohibition, not one conditioned on permission. Independently, every Advance host answers HTTP 403 through DataDome ('Please enable JS and disable any ad blocker', geo.captcha-delivery.com) to the honest user agent on the section page and on the user agreement itself. robots.txt permits /highschoolsports/, which is necessary and not sufficient. FORBIDDEN for every Advance Local site. This paper: /highschoolsports/ returned HTTP 403 with a DataDome interstitial and advancelocal.com/user-agreement/ returned 403 'nginx'; no roundup URL obtained.

  • freep.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: robots.txt permits /story/sports/high-school/ (918-line file, the rest per-bot blocks); the section URL /sports/high-school returned 404 to us, so no live roundup URL was confirmed — stories follow the Gannett pattern via /news-sitemap.xml.

  • detroitnews.com

    Refusedchecked 2026-09-24

    The exception to the MediaNews chain entry: two layers of contract rather than one. The site's own terms (cm.detroitnews.com/terms/) say 'Unless otherwise specified, the Service is intended for your personal, noncommercial use only. You may not modify, copy, reproduce, republish, upload, post, transmit or distribute in any way any material, including code and software, from the Service.' — flat, with no permission route — and the owner's MediaNews Group terms forbid 'robots, spiders, crawlers, data mining or extraction technology ... without written permission'. robots.txt permits and the weekly statewide list (/story/sports/high-school/2026/09/18/michigan-high-school-football-scoreboard-week-4/91826383007/) is served in full behind a soft meter, which changes nothing about the contract. FORBIDDEN as assessed; MediaNews' written permission is the one route named.

  • lansingstatejournal.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: robots.txt permits (identical Gannett template); the section URL /sports/high-school returned 404, so no live roundup was confirmed — 'Greater Lansing high school football Week N schedule, scores' posts follow the Gannett pattern via /news-sitemap.xml.

  • indystar.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: served the section page (/sports/high-school, metered); the statewide Friday 'Indiana high school football scores' list follows /story/sports/high-school/YYYY/MM/DD/<slug>/<id>/ and is listed in /news-sitemap.xml.

  • nwitimes.com

    Refusedchecked 2026-09-24

    Chain terms (/terms/ on every Lee paper; read from tulsaworld.com and journalnow.com): you agree not to 'Access or attempt to access the Site except as expressly permitted in these Terms of Use; ... Use automated scripts to collect information from, or otherwise interact with, the Site', and 'This Site and Our Content may not be copied, reproduced, republished, uploaded, posted, sold, leased, licensed, sublicensed, transmitted, or distributed without our written permission, except that you may download, display, and print one copy of Our Content on a single computer for your personal, non-commercial use only.' robots.txt on the TownNews/BLOX template disallows only CMS plumbing, so the crawl is technically permitted and contractually forbidden — the MaxPreps shape. No permissions route is named, and the platform redirects the honest user agent's second request to a 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). FORBIDDEN for every Lee paper. This paper: the live Friday thread ('Final | Updates and scores from Week 3 of Northwest Indiana high school football', /sports/high-school/football/article_<id>.html) was redirected: the host redirected the honest user agent to a TownNews 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign).

  • journalgazette.net

    Refusedchecked 2026-09-24

    robots.txt permits (TownNews plumbing only). The first request — the standings and schedule article under /sports/high-schools/ — was served in full behind a BLOX meter; the second, for /site/terms.html, was redirected: the host redirected the honest user agent to a TownNews 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). So the site rate-gates this agent after one page. An indexed snippet of the terms, unverified, has users agree 'not to access any of the Content through any automated means (including, but not limited to, use of scripts, web crawlers or screen scrapers) and not to use for commercial purposes or resell any of the data derived from this Content unless specifically allowed in a separate written agreement'. Gate plus an unread prohibition: not assessed to a yes.

  • southbendtribune.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: served the section page; live Friday scoreboard at /story/sports/high-school/2026/09/18/live-south-bend-area-ihsaa-mhsaa-football-scores-updates-sept-18/91786547007/, listed in /news-sitemap.xml. Metered.

  • stltoday.com

    Refusedchecked 2026-09-24

    Chain terms (/terms/ on every Lee paper; read from tulsaworld.com and journalnow.com): you agree not to 'Access or attempt to access the Site except as expressly permitted in these Terms of Use; ... Use automated scripts to collect information from, or otherwise interact with, the Site', and 'This Site and Our Content may not be copied, reproduced, republished, uploaded, posted, sold, leased, licensed, sublicensed, transmitted, or distributed without our written permission, except that you may download, display, and print one copy of Our Content on a single computer for your personal, non-commercial use only.' robots.txt on the TownNews/BLOX template disallows only CMS plumbing, so the crawl is technically permitted and contractually forbidden — the MaxPreps shape. No permissions route is named, and the platform redirects the honest user agent's second request to a 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). FORBIDDEN for every Lee paper. This paper: the weekly collection ('How St. Louis-area football teams fared in Week 3', /sports/high-school/football/collection_<id>.html) was redirected: the host redirected the honest user agent to a TownNews 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign); lee.net/terms/ redirected to a captcha too.

  • kansascity.com

    Refusedchecked 2026-09-24

    Chain terms (mcclatchy.com/terms-of-service, effective 2026-04-14): '10. Use or launch any automated means including spiders, robots, crawlers, scrapers and the like, to download or copy data or content from the Platforms.'; '4.2 ... solely for your personal, non-commercial use ... you will not store or archive a significant portion of the Content or create a database using the Content'; '4.3 Reuse & Republication of Content: You will not reuse, republish or otherwise distribute the Content ... without the express written permission of McClatchy'. The named unlock, mcclatchyreprints.com, is a paid reprint desk rather than a data arrangement, so this is FORBIDDEN rather than LICENSABLE. Independently, every McClatchy host resets the connection for the honest user agent before a byte is served — 'curl: (92) HTTP/2 stream 1 was not closed cleanly: INTERNAL_ERROR' twice, then a 40-45 second HTTP/1.1 timeout with 0 bytes — so robots.txt itself could not be read. Twice refused. This paper: the Akamai edge reset the request for /robots.txt three times (HTTP/2 INTERNAL_ERROR; HTTP:000 over HTTP/1.1); no robots.txt could be read, so no path is known to be permitted.

  • news-leader.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: served the section page; weekly live scoreboard at /story/sports/high-school/2026/09/18/live-scoreboard-for-week-4-southwest-missouri-high-school-football/91754436007/ and a Week 2 MSHSAA scores post, listed in /news-sitemap.xml. Metered.

  • joplinglobe.com

    Refusedchecked 2026-09-24

    robots.txt permits the sports paths, but the first request to the plain user agent (/sports/local_sports/) was redirected: the host redirected the honest user agent to a TownNews 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). A challenge is a gate: stop. Terms (/site/terms.html, CNHI standard) unread — not assessed to a yes. The same platform behaviour on every CNHI paper checked (normantranscript.com, enidnews.com).

  • oklahoman.com

    Not yet askedchecked 2026-09-24

    Chain terms (cm.usatoday.com/terms/): 'you may not ... use robots, spiders, scripts, code or any other automatic or manual device, tool or process to "crawl", "scrape", search or monitor this Site and/or access, retrieve or copy Content or related information'; 'The Site and all its Content is provided solely for your personal non-commercial use'; and, in so many words, 'our robots.txt notice on our Site is intended for designated contractual partners, and it does not constitute our authorization or consent under these Terms of Service.' That last sentence is this register's whole posture written by the other side: a robots allow is not permission. The same terms name the door — 'For information about requesting permission to reproduce or distribute Content from the Site, please contact us' — and describe robots.txt as being for 'designated contractual partners', so the chain is LICENSABLE where its host served the honest user agent: a permissions desk exists and nobody has asked (four of ten assessors read it as forbidden, six as licensable; the enum's own rule is that a clause conditioned on permission is a request unmade, not a refusal). A Gannett host that refused the client outright is FORBIDDEN on its own account. Nothing is fetched until a written yes. This paper: served the section page (/sports/high-school, metered), which at fetch time linked only team-of-the-week polls; no statewide score list confirmed. Stories follow the Gannett pattern via /news-sitemap.xml.

  • tulsaworld.com

    Refusedchecked 2026-09-24

    Chain terms (/terms/ on every Lee paper; read from tulsaworld.com and journalnow.com): you agree not to 'Access or attempt to access the Site except as expressly permitted in these Terms of Use; ... Use automated scripts to collect information from, or otherwise interact with, the Site', and 'This Site and Our Content may not be copied, reproduced, republished, uploaded, posted, sold, leased, licensed, sublicensed, transmitted, or distributed without our written permission, except that you may download, display, and print one copy of Our Content on a single computer for your personal, non-commercial use only.' robots.txt on the TownNews/BLOX template disallows only CMS plumbing, so the crawl is technically permitted and contractually forbidden — the MaxPreps shape. No permissions route is named, and the platform redirects the honest user agent's second request to a 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). FORBIDDEN for every Lee paper. This paper: served the terms (this is where the Lee clause was read) and the 'Week 2 scoreboard for high school football' article — whose body is 138 '<div class="subscriber-only encrypted-content lee-article-text" style="display:none">' blocks with a single visible preview paragraph. Reading it would mean reading behind a paywall, which the posture forbids outright. Forbidden twice over.

  • normantranscript.com

    Refusedchecked 2026-09-24

    robots.txt permits, but the first content request (/sports/) was redirected: the host redirected the honest user agent to a TownNews 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). A gate is an answer: stop. Terms unread — not assessed to a yes.

  • enidnews.com

    Refusedchecked 2026-09-24

    robots.txt permits, but the first content request (/sports/football/) was redirected: the host redirected the honest user agent to a TownNews 'Security Check' captcha (/_services/v1/client_captcha/challenge, service _lb_rate_cmsapp_foreign). Gate: stop. Terms unread — not assessed to a yes.

  • seattletimes.com

    Refusedchecked 2026-09-24

    robots.txt permits the path but its own header says 'Use of any device, tool, or process designed to data mine or scrape the content using automated means is prohibited without prior written permission from The Seattle Times Company', and Section 20 of the terms (/notices/terms.html): 'You agree that you will not use any robot, spider, scraper or other automated means to access the Sites for any purpose without our express written permission.' Moreover the scoreboard page (/high-school-football-scoreboard/) is a single '<iframe src="https://scorebooklive.com/widgets/v3/73?group_id=69">': the scores are SBLive's widget, and sblive is already FORBIDDEN in this register — the same door with a different sign.

  • spokesman.com

    Refusedchecked 2026-09-24

    robots.txt permits /stories/ and there is no bot clause, but the Service Agreement (/service-agreement/) licenses the Contents for 'personal, noncommercial use' only: 'You may not modify, publish, transmit, participate in the transfer or sale of, reproduce ..., create new works from, distribute, perform, display, or in any way exploit, any of the Contents' and 'Copying or storing of any Contents for other than personal use is expressly prohibited without prior written permission from us or from the copyright holder'. Storing Friday's results in a commercial product is copying and storing for other than personal use, so the contract forbids it as written even though a score is a fact. The clause names the remedy, 'prior written permission'; an ask-first target for Eastern Washington and North Idaho, forbidden until granted. Weekly roundup at /stories/YYYY/mon/DD/<slug>/, metered.

  • heraldnet.com

    Refusedchecked 2026-09-24

    robots.txt is fully permissive ('User-agent: * / Disallow:') and there is no automated-access clause, but the Carpenter Media Group terms (carpentermediagroup.com/terms-of-service/, linked from the footer) say 'You may not modify, copy, reproduce, republish, upload, post, transmit or distribute in any way any material from the Service' and 'You may download material from the Service for your personal, non-commercial use only'. Hometown Line's use is commercial republication of results, so the licence does not cover it as written. With an open robots.txt, a soft meter and a roundup that even solicits scores from coaches, the best ask-first target in Washington; forbidden until Carpenter/Sound Publishing says yes in writing. Roundup at /YYYY/MM/DD/<slug>/ (body opens 'Prep football roundup for ...'), via /sitemap-news.xml.

  • thenewstribune.com

    Refusedchecked 2026-09-24

    Chain terms (mcclatchy.com/terms-of-service, effective 2026-04-14): '10. Use or launch any automated means including spiders, robots, crawlers, scrapers and the like, to download or copy data or content from the Platforms.'; '4.2 ... solely for your personal, non-commercial use ... you will not store or archive a significant portion of the Content or create a database using the Content'; '4.3 Reuse & Republication of Content: You will not reuse, republish or otherwise distribute the Content ... without the express written permission of McClatchy'. The named unlock, mcclatchyreprints.com, is a paid reprint desk rather than a data arrangement, so this is FORBIDDEN rather than LICENSABLE. Independently, every McClatchy host resets the connection for the honest user agent before a byte is served — 'curl: (92) HTTP/2 stream 1 was not closed cleanly: INTERNAL_ERROR' twice, then a 40-45 second HTTP/1.1 timeout with 0 bytes — so robots.txt itself could not be read. Twice refused. This paper: /robots.txt reset the connection for the honest user agent over both HTTP/2 and HTTP/1.1 (same Akamai behaviour as kansascity.com); no robots.txt could be read.

Nothing is fetched from a site whose terms or robots file refuse it, and no block is worked around. How the rest of the site keeps these distinctions is on the compliance page, and why most of the season had to be generated is on the about page.