Origenality

Method

Origenality is a bibliographic map of the scholarship on Origen of Alexandria. It gathers the records of open catalogues, keeps one record per work, files each under a controlled vocabulary of themes, works of Origen and approaches, and draws the result so that a reader sees where the field is thick and where it is thin.

This page says how the corpus is assembled, what is published and under what terms, and what the density figure does and does not measure.

Version of August 2026 · one source · 1 632 records harvested, 1 402 counted

1. The corpus

One catalogue today, named as such everywhere on the site.

The corpus behind this version is a single harvest: the Index Theologicus (IxTheo, Tübingen), taken through the authority record for Origen and hydrated from K10plus. On 15 August 2026 that harvest returned 2 116 records. Of these, 484 are editions and translations of Origen's own works — primary sources, set aside — and 1 632 are publications about him. That second figure is the corpus. Adding the 47 records that are both an edition and a collection of studies to it would count them twice.

Every record keeps the catalogue number it was harvested with, and every entry in the Explorer links back to the record it came from. Nothing on this site replaces a catalogue; it points at one.

What is not here yet

Nine further harvests exist in the working data — OpenAlex, Crossref, Semantic Scholar, Isidore, theses.fr, Dialnet, the Italian SBN, the Adamantius repertorio, the BIBP of Laval. No record of theirs is on this map, and no figure on this site counts one. Their relevance filtering is not done, and a figure drawn from an unsorted pile would be worse than no figure. The site says one source because its corpus is one source.

They serve one purpose here, and it does not touch the counts. Where one of them summarises a publication that IxTheo also catalogues, that summary is attached to the IxTheo record and credited to the database that wrote it (§ 2). It adds prose to a record already in the corpus; it adds no record, and it moves no figure.

A count from a single catalogue measures that catalogue. IxTheo indexes German and European theology closely; its subject headings are richer on recent records. Both facts push the counts, and both are visible in the Observatory.

2. Rights, and what is published

Attribution on every summary, and removal for the asking.

An author, a title, a year, a journal, an identifier: bibliographic metadata are facts, and K10plus publishes them under CC0. They are what this site republishes: title, authors, year, language, container, publisher, DOI, ISBN, subject headings, format. A field is published where the record carries it, and nothing stands in its place where it does not: 476 of the 1 632 records name a publisher, 358 an ISBN, 130 a DOI.

Abstracts are shown as well. A reader who sees only a title cannot tell whether a study answers the question they came with, and a bibliographic map that withholds the one paragraph that would settle it serves nobody. So every summary in the Explorer is displayed, and every summary displayed names the database that wrote it and links to the record there. A summary written for this project says so instead of naming a database.

Removal on request. A publisher, a partner database or an author who asks for their summaries to be taken down gets it, without argument and without a negotiation. One message to romain.girardi@univ-cotedazur.fr naming the database or the imprint is enough; the summaries come off the site and out of the next data release. The partner databases were told in writing on 15 August 2026, before any of this was published. The procedure is repeated on the Credits page, next to the licence of each source.

What the source declared about its own rights is kept with each record, internally, so that a request can be honoured in one pass rather than by hand. It is a trace, not a gate: it decides nothing about what is shown.

Coverage. IxTheo is a catalogue, and a catalogue describes rather than summarises: 152 of the 1 632 records carry a summary of their own. Where another database summarises the same publication, that summary is joined to the record by catalogue number, by DOI, or by title and year, and 229 records are covered that way. In all, 381 records of 1 632 (23.3 %) show a summary. The other 1 251 are not being withheld: 1 212 of them appear in the working corpus and no database there holds a summary of them either.

Who wrote them. How a summary reached this file is one question; who wrote it is another, and the second is the one the credit answers. 166 of the 381 were written by the catalogue itself and 215 by another database, whichever route they took to get here. The full breakdown, counted from the published file rather than typed here:

Database that wrote the summarySummaries
Index Theologicus (IxTheo / K10plus)166
OpenAlex112
GIROTA / Adamantius89
ISIDORE (Huma-Num)6
Crossref5
Semantic Scholar3

3. The semantic vocabulary

Closed lists with stable identifiers, not coordinates in a latent space.

Each record is described on four independent axes, each a versioned file of closed values:

AxisSizeWhat it holds
Works33 the works of Origen by their usual Latin titles and sigla, plus the default value for a study that names none
Themes16 / 61 sixteen domains, sixty-one leaves; only a leaf is ever written on a record, the domain is for aggregation and for the map
Approaches10 philological, exegetical, doctrinal, historical, reception, comparative, methodological, edition, review, survey
Relevance4 the four classes that decide whether a record counts in a density figure

The theme axis was built by crossing two sources, and the file keeps the trace of both: the thirty-seven published sections of the Adamantius repertorio, which give the axis of author and milieu, and the subject headings of the harvest itself, which give the doctrinal and exegetical axis. Every leaf lists the headings that motivated it.

Labels are given in English, German, French and Italian. The German label is the catalogue heading itself wherever one exists, the Italian follows the wording of the repertorio, and any label supplied by the vocabulary rather than found in a source is marked as supplied. The Explorer shows the English label on the map and the other three on the hover card, so that a reader who knows the field under its German headings can still recognise the cluster.

How a record is tagged

The tagging program sends the metadata of one record against the four vocabularies and accepts only values that exist in them: the schema is derived from the vocabulary files, so it cannot drift from them. A value that falls outside is dropped at validation, the repair is written into the record, and the record is flagged for review. Records that fail to be read at all are kept in a rejects file with their cause; none disappears. Identifiers are deterministic, so two runs address the same objects, and each record carries the vocabulary version, the prompt version and a digest of the exact input submitted. Changing the vocabulary means a new version and a new wave, so that tags made under different instructions stay comparable. A correction inside one wave is appended rather than overwritten, and the file is compacted at the end of the wave, its superseded lines kept in a history beside it. One wave file has been damaged by a renumbering tool and rebuilt from the site asset, losing its confidence and justification fields; it is kept as an archive, and nothing reads it.

The hand-tagged reference set, and where it stands

The project holds itself to one figure: on a set of records tagged by hand, the program and the hand must agree at least nine times in ten on the question that decides every count — does this record enter the density or not. The set was fifty records for the second wave, drawn with a fixed seed and tagged against the same vocabulary and the same instructions. The tags this version draws come from that second wave, run over the whole working corpus and joined back to the catalogue records this map shows.

No current agreement figure is published on this page. The previous measurement, taken under instructions that have since been rewritten, came in under the bar. Six records decided it, and they fall on two boundaries the earlier instructions had left to judgement — an old printed edition of a text of Origen, and the line between a study where he holds an identifiable section and one where a cataloguer's classification is all there is. Both are now settled in writing, and the second is doubled by a rule in the validator that no answer can override.

The fifty have since been tagged again by hand under those settled instructions, and the figure will be published from a run under the same instructions and from nothing else. Comparing the new hand set against the second wave — the run whose tags this version draws — gives 48 of 50 on the density question, but that comparison moves the reference rather than the program, and it is not the measurement: it is recorded in the working notes, not printed as a validation. Republishing an earlier figure as a current validation would be reporting a test that was never run under the present instructions.

A known bias, and what became of it. The rule used in the first wave demanded that the metadata name Origen positively, and filed everything else as not about Origen. On a curated perimeter, where cataloguers have already attached each record to Origen's authority record, that rule was too severe: a study of original sin in the Fathers, or of ministry in the early Church, was filed out although Origen is one of its witnesses. Seven of the eight disagreements in the thirty-record pilot pointed that way. The instructions now floor the class at mentioned only for any record catalogued as being about him, and keep not about Origen for homonyms and for texts by Origen catalogued as texts about him. The harvest was classed again under them: the reservoir holds 5 records where it held 223, and mentioned only holds 225 where it held 9 — retrieved by a search, counted in no figure. What remains of the bias runs the other way now: a record the catalogue attached to Origen is credited with a mention even where the metadata alone would not carry it.

4. The density figure

A compass toward thin ground, not a verdict on originality.

What it is. A count of records inside a named node of the vocabulary: 47 works in this perimeter means forty-seven records of this harvest are filed under that node. The node has a name, a path and a definition, so the figure can be checked. It always travels with four things: the source corpus, the wave that produced the tags, the relevance threshold applied, and the share of the node flagged for review.

What it is not. It is not a measure of the originality of a project, and not a measure of quality. A thin area may be thin because nobody has looked, or because there is nothing to find, or because this catalogue does not index the language in which the work was done. The figure opens a question; it does not close one.

Which records count. Every figure on this site is a count of the same 1 402 records: those classed core or partial, where Origen is the subject or holds an identifiable section of the argument. Of the 1 632 harvested, 225 are classed mentioned only and 5 are held outside the count; neither class enters a figure, on any page — not the bars of the Observatory, not the language counts, not the number on a question chip, not the density of a cluster.

Retrieval is wider than counting, and stays wider. A record where Origen is mentioned only still answers a search and still answers the four questions: nothing is put aside, as the map says on every page. It is returned, listed and readable — and it is not added to the figure. Where a surface can return more than it counts, it says so in words and gives the second number: N further works are mentioned only and are listed below the count.

What the count leaves out, and does not hide. Every one of the 1 632 records carries a class. Twenty-eight did not: their cluster had been split or joined by the deduplication after the wave had run, or a mechanical pre-sort had set them aside, and the count held them outside every figure for want of a tag rather than guess one. They were sent back to the tagger in a later pass, and the build now refuses to publish while a record on display carries no class.

What it does not include. The size of a disc on the map is an academic weight — the mean of a citation percentile inside the work's own cohort (same decade, same document type, same language) and a structural percentile in the graph. That weight is drawn, and only drawn. It never filters, it never removes anything from a search, and it never enters the density. A record without a citation figure keeps the base size and reads no citation data rather than being pushed down.

The citation figure is unevenly available, and the unevenness runs along languages. Citation counts come from a single measurable source, and it does not cover the field evenly. Across the working corpus of 42 210 clusters, 35.0 % carry a count at all, and the share by language of publication is Spanish 88.9 %, German 39.8 %, English 35.9 %, Italian 24.4 %, French 9.6 %. A French article is not less cited than a Spanish one; it is less often indexed where citations are counted. This is why the weight is a cohort rank rather than a raw count — a record is ranked against others of its own decade, type and language, so that a thinly indexed language is not read as a thinly cited one — and why the weight never filters anything out of a result. Read across languages, the disc sizes still carry that bias, and no correction here removes it.

No reading or download figures appear anywhere on this site. No source holds them reliably, and simulating one would be an invention.

5. Nothing is thrown away

The map has reservoirs rather than a bin. Records classed as not about Origen have one; publications that bear on no single work of Origen have theirs. A reservoir is folded under the field, named and counted, and it is drawn on the map when the reader asks for it: the named clusters are what the field is for, and a reservoir left open would be the widest cloud on it. They remain searchable, and they remain in every export.

One rule for the counts, on this page, in the questions of the Explorer and in the Observatory alike: a figure counts the 1 402 records classed core or partial (§ 4), and nothing else. The reservoirs hold the rest, named and counted as reservoirs. Saying how much of a harvest cannot be placed is part of what a map of a field owes its reader.

6. Reproducibility

  1. The vocabulary is versioned and cited on every record.
  2. The instructions given to the classifier are versioned and can be printed word for word.
  3. The schema is derived from the vocabulary, not copied from it.
  4. The exact input submitted for each record is fingerprinted.
  5. The hand-tagged reference set fixes the human reference and is published with the run, agreement figure and disagreements alike.
  6. Rejects and review rates are published with any figure drawn from the tags.
  7. Every density figure states its perimeter.

A third party who runs the same vocabulary, the same instructions and the same schema on the same records can measure the difference record by record. That is the sense in which a classification of this kind can be checked at all.

7. Questions

Short answers, each standing on its own.

What does the density figure measure?

A density is a count of records inside a named node of the vocabulary, and nothing more. When the map says forty-seven works sit in a perimeter, forty-seven records of this harvest are filed under that node, out of the 1 402 the site counts. The node has a name, a path and a definition, so the figure can be checked against the same data. What a density does not measure is the originality of a project: thin ground may be unexplored, or exhausted, or written in a language this catalogue indexes poorly, and the map cannot tell those three apart. It opens a question rather than settling one. Academic weight, drawn as the size of a disc, is a separate thing again: a citation percentile inside a cohort of the same decade, type and language, which never filters a search and never enters a count.

Which sources does the map draw on?

One catalogue, and the site names it on every page: the Index Theologicus of Tübingen, hydrated from K10plus and harvested through the authority record for Origen, GND 118590235, on 15 August 2026. Its metadata are published under CC0, and every entry in the Explorer links back to the record it came from. Nine further harvests exist in the working data: OpenAlex, Crossref, Semantic Scholar, Isidore, theses.fr, Dialnet, the Italian SBN, the Adamantius repertorio and the BIBP of Laval. None of them puts a record on this map and none of them moves a figure, because their relevance filtering is not finished. They serve one purpose here: where one of them summarises a publication the catalogue also holds, that summary is attached to the record and credited to whoever wrote it. Subscription indexes cannot be harvested and are absent, so coverage before 1980 is thinner than the counts suggest.

How are the summaries credited, and how are they removed?

Every summary shown in the Explorer names the database that wrote it and links to the record there; a summary written for this project says so instead of naming a database. The regime is attribution and takedown. A publisher, a partner database or an author who wants their summaries off the site writes to romain.girardi@univ-cotedazur.fr, naming the database, the imprint or the article, and those summaries come off the site and out of the next data release. No justification is asked for and none is needed. The removal is scripted rather than done by hand, so it lands in one pass and cannot be half applied. The record keeps its bibliographic metadata and its link, so nothing is lost to the reader but the paragraph. The partner databases were written to on 15 August 2026, before any of this was published.

How should Origenality be cited?

In a footnote: Romain Girardi, Origenality: a bibliographic map of Origen studies, 2026, https://origenality.com. The machine-readable form is in CITATION.cff at the root of the repository, which carries the version, the date of release and the affiliation. Cite the underlying databases as well when you cite a record: each source has its attribution template in ATTRIBUTION.md, and Semantic Scholar in particular requires attribution under ODC-BY. The site is a pointer rather than a substitute, and a record found here should be read in the catalogue that describes it. The code is under the MIT licence; the data are published under attribution and takedown, source by source, as DATA_POLICY.md sets out, and no single licence is asserted over the whole.

Origenality · Romain Girardi Observatory Credits and licences Method, August 2026