How to Scrape and Extract Hyperlink Networks with BeautifulSoup and NetworkX

Our great sponsors

InfluxDB - Power Real-Time Data Analytics at Scale

WorkOS - The modern identity platform for B2B SaaS

SaaSHub - Software Alternatives and Reviews

Our great sponsors

grub-2.0

4 19 0.0 Python

Grub is an AI powered Web crawler.

Depending on the use case you might try imaging the page, then send the image to an ML model for full text before indexing. If you need links extracted, Selenium also supports parsing the assembled DOM: https://github.com/kordless/grub-2.0/tree/main/aperture
argos-search

1 6 0.0 Python

Small search engine

I wrote a similar Python library to do Beautiful Soup scraping with basic PageRank and a Flask web app: https://github.com/argosopentech/argos-search
InfluxDB

www.influxdata.com
sponsored

Power Real-Time Data Analytics at Scale. Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.

NOTE: The number of mentions on this list indicates mentions on common posts plus user suggested alternatives. Hence, a higher number means a more popular project.

Suggest a related project