About a decade ago I tried to create an alternative to Google, using the Dewey Decimal System to group web domains by subject matter. My vision was that the website was a library, and domains were individual books. Rankings were a mixture of domain reputation and cultural significance.
The OCLC is very protective of their IP, and likes to sue.
Needless to say, my project was shut down before even getting off the ground.
I’m working on reviving the project, under a new classification system (The Open Categorization System), and would love any help I can get.
The project can be found here: https://github.com/ki4jgt/Open-Catalog-System
I’m curious as to why you’d limit this to the domain level. A university for example might have thousands of URLs in it with wildly different subjects:
university.tld/math/student-name/thesis-on-mathy-subject/university.tld/journalism/student-name/big-story-about-politics
How would your system account for this?
I’m guessing that people would just find the content with the site’s own navigation / search mechanism.
I don’t really see how it would make sense to replicate URLs below a domain level. Or how it would be manageable.
and I don’t see any value in limiting classification to the domain level. How would one classify
wikipedia.orgin this scenario? Would it not make more sense to define an open standard that’d leverage this system but allow domain managers to define it themselves?To take my
university.tldas the example, that university might host a file called athttps://university.tld/ocs.jsonthat looks something like this:{ "/": "EDU" "/math": "MAT", "/journalism": "JOU", }(Heads up to OP: there’s no journalism in your current spec. That feels like an oversight.)
This would allow the high-content site admins to classify parts of the site differently. The
/ocs.jsonfile might even be dynamically generated for complex sites like Wikipedia.
What sort of help are you looking for?



