About a decade ago I tried to create an alternative to Google, using the Dewey Decimal System to group web domains by subject matter. My vision was that the website was a library, and domains were individual books. Rankings were a mixture of domain reputation and cultural significance.

The OCLC is very protective of their IP, and likes to sue.

Needless to say, my project was shut down before even getting off the ground.

I’m working on reviving the project, under a new classification system (The Open Categorization System), and would love any help I can get.

The project can be found here: https://github.com/ki4jgt/Open-Catalog-System

  • Daniel Quinn@lemmy.ca
    link
    fedilink
    English
    arrow-up
    3
    ·
    6 hours ago

    I’m curious as to why you’d limit this to the domain level. A university for example might have thousands of URLs in it with wildly different subjects:

    • university.tld/math/student-name/thesis-on-mathy-subject/
    • university.tld/journalism/student-name/big-story-about-politics

    How would your system account for this?

    • Radieschen@slrpnk.net
      link
      fedilink
      arrow-up
      1
      ·
      6 hours ago

      I’m guessing that people would just find the content with the site’s own navigation / search mechanism.

      I don’t really see how it would make sense to replicate URLs below a domain level. Or how it would be manageable.

      • Daniel Quinn@lemmy.ca
        link
        fedilink
        English
        arrow-up
        1
        ·
        edit-2
        1 hour ago

        and I don’t see any value in limiting classification to the domain level. How would one classify wikipedia.org in this scenario? Would it not make more sense to define an open standard that’d leverage this system but allow domain managers to define it themselves?

        To take my university.tld as the example, that university might host a file called at https://university.tld/ocs.json that looks something like this:

        {
          "/": "EDU"
          "/math": "MAT",
          "/journalism": "JOU",
        }
        

        (Heads up to OP: there’s no journalism in your current spec. That feels like an oversight.)

        This would allow the high-content site admins to classify parts of the site differently. The /ocs.json file might even be dynamically generated for complex sites like Wikipedia.