trafilatura
docker-languagetool
trafilatura | docker-languagetool | |
---|---|---|
13 | 10 | |
2,853 | 402 | |
- | - | |
8.7 | 5.9 | |
2 days ago | about 2 months ago | |
Python | Shell | |
Apache License 2.0 | GNU Lesser General Public License v3.0 only |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
trafilatura
-
Trafilatura: Python tool to gather text on the Web
The feature list answers that question pretty well: https://github.com/adbar/trafilatura#features
Basically: you could implement all of this on top of BeautifulSoup - polite crawling policies, sitemap and feed parsing, URL de-duplication, parallel processing, download queues, heuristics for extracting just the main article content, metadata extraction, language detection... but it would require writing an enormous amount of extra code.
-
Show HN: Build AI Dags with Memory; Run and Validate LLM Tools in Containers
The WebScraper tool uses Trafilatura [1] to scrape and parse HTML—nothing too fancy. "Scraping" a React site would require a totally different approach, probably something more akin to Adept's ACT-1 [2].
I run a local chat app built with Griptape and I use it to give me summaries of web pages or answer specific questions all the time :)
1. https://github.com/adbar/trafilatura/
-
Powerful and free scraper with a headless browser under the hood and Readability for parsing
I've been playing with Trafilatura lately, and it's very good. There are a few very thorough comparisons to other projects and it really shines. It doesn't do anything headless from what I can tell, but it doesn't have to do the scraping itself. Maybe an option could be to use Playwright to scrape, then Trafilatura to parse. Food for thought.
-
I made a Chrome Extension that lets you ask any question about the page you are on (bluf.ai)
Cool! If you care to explain me further... :) ... I tried parsing a page using: https://github.com/adbar/trafilatura, json stringify it and passing it to https://platform.openai.com/docs/api-reference/embeddings/create. How do I use the response as an input later? <3
-
Testing fast installation in tear-down environment
I want to test how easy it is to install a package plus special extra dependencies to run a certain script in that package: https://github.com/adbar/trafilatura
- Advice on standard design pattern for comparison test script
- Automate dependency installation
- Issue with sklearn
- Questions about some code
- How does Firefox's Reader View work?
docker-languagetool
- LanguageTool and Plagiarism
-
Is there an open-source alternative for Grammarly with a free proprietary license?
LanguageTool itself is open source (basic features). I'm running it in a Docker container.
- Is ProWritingAid the same as Grammarly when it comes to security?
-
Ask HN: Do You Trust Grammarly?
For developers who already have Docker running on their machine. I can strongly recommend running it locally with e.g. Docker Compose.
Safes effort with maintaining an installation and keeping the background process running. Plus, it also works when network connectivity drops.
https://github.com/Erikvl87/docker-languagetool/blob/master/...
-
Setup your private LanguageTool server
Following, you can see my personal docker-compose.yml, which you can use as a reference. For a more detailed description, you can look at the image description erikvl87/languagetool.
-
What do we say to typos? Not today!
I was already used to wiggly lines in my favorite IDE IntelliJ and really missed the spell and grammar check capabilities in other editors especially when writing something in the browser. A colleague told me that IntelliJ is using LanguageTool since I'm pretty satisfied with the analysis inside it. Therefore, I looked around on GitHub for a way of hosting my own LanguageTool server. I came across this repository and decided to give it a go and run it on my Linux server.
-
LanguageTool – FOSS Style and Grammar Checker for 25 Languages
It's great, I've been running the self-hosted version for a few months now. I run it using a container[0] on a $10/mo Digital Ocean VPS using Dokku[1]. The main downside is that it's a bit of a memory hog, possible because it's written in Java? Otherwise, I haven't had to mess with it after the initial set up. I like that I'm not sending everything I write to another 3rd party service I don't control.
As for the plugin, it definitely catches more issues than the stock browser spellchecks. The main issue I have with it (maybe someone can point me in the right direction) is that it always tries to autodetect the language. This is fine for longer texts, but often fails on shorter strings like headlines. This leaves my words highlighted red because it thinks I'm writing bad German, or Swedish (which is fair). I haven't been able to figure out how to force it to only use US English.
[0]: https://hub.docker.com/r/erikvl87/languagetool
-
What's something self hosted everyone needs to run ?
https://github.com/Erikvl87/docker-languagetool became my daily tool
-
Language Tool - Grammarly Alternative
The entire proof reading engine can be run either through their Java server or in a docker image. Java Docker Another Docker Repo
-
Anyone self-hosting languagetool?
I have been using this docker version without any issues: https://github.com/Erikvl87/docker-languagetool
What are some alternatives?
newspaper - newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:
languagetool - Style and Grammar Checker for 25+ Languages
python-goose - Html Content / Article Extractor, web scrapping lib in Python
Readarr - Book Manager and Automation (Sonarr for Ebooks)
TWINT - An advanced Twitter scraping & OSINT tool written in Python that doesn't use Twitter's API, allowing you to scrape a user's followers, following, Tweets and more while evading most API limitations.
project-zomboid - A Project Zomboid server with LinuxGSM.
html2text - Convert HTML to Markdown-formatted text.
Invidious - Invidious is an alternative front-end to YouTube
Goose3 - A Python 3 compatible version of goose http://goose3.readthedocs.io/en/latest/index.html
nginx-rtmp-docker - Docker image with Nginx using the nginx-rtmp-module module for live multimedia (video) streaming.
textract - extract text from any document. no muss. no fuss.
whoogle-search - A self-hosted, ad-free, privacy-respecting metasearch engine