Back to Internet Archive
GSoC 2026

From Messy Subjects to First-Class Genre Tags

Open Library has millions of book records tagged with inconsistent, unstructured subject strings, for example, "science fiction", "Science Fiction", and "sci-fi" are treated as completely different tags, causing patrons to miss thousands of relevant books depending on how they search. This project solves that by building system-wide support for first-class genre tags on Works, directly addressing issue #11610. The solution works in four connected layers: first, developing a controlled genre vocabulary and a mapping dictionary from messy subject strings to canonical genre labels; second, updating the work schema in infogami to add a genres field storing canonical strings that serve as keys to fetch corresponding Tag objects; third, building a batch backfill pipeline that processes tens of thousands of high-demand works and populates their genres field; and fourth, updating Solr to index and facet by genre, displaying genre chips on book pages, and adding a librarian editing UI with autocomplete. By the end of GSoC, a patron searching for science fiction will get consistent, accurate results regardless of how they type it because works will have clean genre tags, indexed by Solr, backed by canonical Tag objects, and maintainable by librarians.

Project details

Contributor

Chisom Nnamani

Mentors

Not available

Technologies

Not listed in the archive