Clarrel Commons — Terms

Terms version: v1

What Clarrel Commons is

Clarrel Commons is a public, non-commercial digital library of primary philosophical texts and their translations, built for search and scholarly research. It is provided free of charge and will remain so — any future monetization is a decision that, if it ever happens, applies outside this tier, never to the Commons itself.

The corpus draws on many sources, under different licences

The texts in Clarrel Commons come from a range of sources — public philological repositories, public-domain digitizations, and library archives — and they do not all carry the same licence. Some examples, so this isn't an abstract claim:

Because of this mix, there is no single blanket licence for "the corpus." Each source's specific licence and rights basis is tracked as part of its own record. Where a licence requires attribution, that obligation travels with whatever specific text you're looking at, not with the corpus as a whole.

Attribution and ShareAlike

Where a source's licence requires attribution or ShareAlike (CC BY-SA content, described above), that obligation applies to the original source text as redistributed, including anything you derive or export from it.

What we collect when you search

When you search, a few anonymous, aggregate usage signals may be recorded -- never anything that identifies you. None of them link your searches to each other. One of them (feedback button clicks, below) links to other feedback from the same visit only, using a temporary code that is never linked to a search history, an account, or an identity, and never survives past that visit: - The item IDs of the passages that were returned -- which passages the Commons corpus surfaces for search queries, nothing more. - Your search terms, but only after named entities (such as email addresses or phone numbers) are removed, rare terms are suppressed, and the remaining terms are counted in aggregate across a batch of many searches -- never your individual search text, and never linked back to you or to any other search. - A rough daily count of distinct visitors, produced from a one-way code derived from your connection that is regenerated every day. The code is never stored and cannot be linked across days or back to an address. This count is necessarily imprecise: several people sharing one network connection can undercount as a single visitor, and one visitor whose address changes (for example, over a VPN) can overcount as several. - If you click a "more relevant" button, which passage you marked as more relevant than the one above it -- never your search text, and never anything else about you. - If you click a "useful" or "not useful" button, which passage you marked and that verdict, together with a one-way code from a temporary cookie this site sets for that purpose. The cookie holds no information about you, is not shared with anyone, and is automatically cleared when you close your browser -- it exists only so multiple pieces of feedback from the same visit can be linked to each other, never across separate visits, and never to a search history, an account, or an identity. This is the only cookie this site sets that relates to the usage signals described above. A separate, long-lived cookie records that you've acknowledged this site's terms (see the terms page) so you aren't asked again on your next visit -- unrelated to any of the above, and also never linked to a search history, an account, or an identity. This is separate from ordinary web server logging. This service does not log search query content at all, in any log, for any purpose. It does keep one operational log, with no request content of any kind in it: a one-way code derived from the requesting connection that changes every day -- the same kind of code used for the daily visitor count above, so that log can never be linked back to an address -- together with which page or endpoint was requested and its status, so that abuse (for example, one connection making an unreasonable number of requests) can be noticed and addressed. It is not analyzed for individual behavior otherwise, and is kept for 14 days and then automatically deleted.

Rate limits and bulk access

To keep the service usable for everyone, /search and /api/search are rate-limited per visitor (and, for API callers that declare a client id, per client): requests are currently allowed in bursts of up to 10, sustained at up to 20 per minute. Exceeding these limits returns 429 Too Many Requests with a Retry-After header rather than a silent failure. These specific numbers are provisional — measured against a benchmark corpus rather than the Commons' full, live corpus — and may be revised; this page states whatever the current limits actually are.

Beyond ordinary rate limits, reasonable-use limits on bulk enumeration and large-scale scraping of the search surface apply. The specific policy for that is still being worked out (tracked as avophile/clarrel-commons#81) — for now, "reasonable use limits apply" is the operative rule, and this section will be updated with the specific policy once it lands. Sanctioned bulk export for members of a personal Clarrel instance is a separate, authenticated channel outside this public surface, not something reachable through ordinary search.