I am one of a network of academic researchers from around the world working on collecting media market data. One problem is that referenced sources often disappear which makes validation later difficult or impossible. So, I thought I would recommend self-hosting something like archive.org that would allow affiliated researchers to submit their web references and have their sources efficiently archived in a central project repository. That would allow validation and continuity for when web-hosted text and files disappear or researchers leave.

I have been looking at ArchiveBox. If you have experience of this or a similar solution, would that fit the bill? The important thing is efficiency for researchers submitting/retrieving pages and files, and openness in structure and formats so that the archive would remain useful if ArchiveBox or similar disappears. FOSS of course means you can’t be locked out anyway.

  • Stopwatch1986@lemmy.mlOP
    link
    fedilink
    English
    arrow-up
    2
    ·
    6 hours ago

    A wiki is a good idea. Putting a Singlefile or similar all-in-one file in a repository and provide index numbers organised as a look-up table would also work for easy retrieval by a random research user. Both require some admin and more effort from the researchers.

    I wish there was a hostable version of archive.is for near-zero maintenance. You just submit a URL over the internet and the web page is cached once along with a screenshot. Then, anyone can access the archived version. This can be done already with archive.is but we have no control over its future, which is critical for long-term dependable archiving.

    • irmadlad@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      3 hours ago

      This can be done already with archive.is but we have no control

      Did a little digging this morning. I honestly can’t find a selfhosted, archive.is alternative. All the solutions I came up with are either paid for and online use only, or free, but still online use only.