mwparserfromhell
pico
mwparserfromhell | pico | |
---|---|---|
5 | 8 | |
708 | 720 | |
- | 61.9% | |
6.6 | 9.6 | |
17 days ago | 1 day ago | |
Python | Go | |
MIT License | MIT License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
mwparserfromhell
- FLaNK AI Weekly for 29 April 2024
-
Processing Wikipedia Dumps With Python
There's also https://github.com/earwig/mwparserfromhell, if you don't want to roll your own.
-
[Python] How can I clean up Wikipedia's XML backup dump to create dictionaries of commonly used words for multiple languages?
In particular what you're looking at is not XML but wikitext. I found a discussion on stackoverflow about solving the same problem of getting text from wikitext. Seems like the most promising solution in Python since you already have the dump is to run each page through mwparserfromhell. According to the top stackoverflow answer you could use something like
-
How can I clean up Wikipedia's XML backup dump to create dictionaries of commonly used words for multiple languages?
Thank you so much! I was actually talking about the markup language within the text. Turns out it's proprietary to WikiMedia and user lowerthansound kindly suggested I use this: https://github.com/earwig/mwparserfromhell
pico
- FLaNK AI Weekly for 29 April 2024
-
Pico.sh – Hacker Labs
The repo is terrible at tell us what is this about, the landing page is better: https://pico.sh, but still terrible.
-
Show HN: Pgs.sh – A zero-install static site hosting service for hackers
Thanks for the feedback! We deployed a change to support avif: https://github.com/picosh/pico/commit/570514201d926a664c88cb...
-
SSH3: SSH using HTTP/3 and QUIC
SNI is absolutely needed. Over at https://pico.sh we have to request an IP for each ssh server even though from a resource perspective we really only need 1 VM. It increases the complexity of our deployments and overall makes us want to figure out how to merge all of our SSH apps into one.
-
Show HN: Pgs.sh – A zero-dependency static site hosting service for hackers
Yes! We have a monorepo with a bunch of services, but it's all here: https://github.com/picosh/pico
What are some alternatives?
wikitextparser - A Python library to parse MediaWiki WikiText
archwiki - MediaWiki used on Arch Linux websites (read-only mirror)
WiktionaryParser - A Python Wiktionary Parser
wikiteam - Tools for downloading and preserving wikis. We archive wikis, from Wikipedia to tiniest wikis. As of 2023, WikiTeam has preserved more than 350,000 wikis.
pywikibot - A Python library that interfaces with the MediaWiki API. This is a mirror from gerrit.wikimedia.org. Do not submit any patches here. See https://www.mediawiki.org/wiki/Developer_account for contributing.
isbntools - python app/framework for 'all things ISBN' including metadata, descriptions, covers...
wiki_dump - A library that assists in traversing and downloading from Wikimedia Data Dumps and their mirrors.
pastevents - A structured, searchable archive of Wikipedia's "Current Events" portal
wikifunctions - Python functions for retrieving data from the MediaWiki/Wikipedia API
Wiki-scripts - Scripts used on the official factorio wiki
MediaWiki-Tools - Tools for getting data from MediaWiki websites