PdfPig vs NetworkX

Our great sponsors

WorkOS - The modern identity platform for B2B SaaS

InfluxDB - Power Real-Time Data Analytics at Scale

SaaSHub - Software Alternatives and Reviews

Our great sponsors

PdfPig		NetworkX
	Project
7	Mentions	61
1,462	Stars	14,178
4.9%	Growth	1.6%
9.1	Activity	9.6
9 days ago	Latest Commit	6 days ago
C#	Language	Python
Apache License 2.0	License	GNU General Public License v3.0 or later

The number of mentions indicates the total number of mentions that we've tracked plus the number of user suggested alternatives.
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.

PdfPig

Posts with mentions or reviews of PdfPig. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2022-11-30.

Just Say No
3 projects | news.ycombinator.com | 30 Nov 2022

Maybe (most likely) this is a problem of GitHub's terminology. For genuine bugs, e.g. here's the repro, the stack trace, the code to replicate it, it happens 100% of the time if you follow these steps, I'd agree that just having it open and in the backlog would be preferable.
The problem is those make up maybe at a generous estimate, 10-15% of issues in a projects backlog. In the interests of full disclosure here's mine (I don't use stalebot) https://github.com/UglyToad/PdfPig/issues?page=1&q=is%3Aissu.... As you can see from the backlog I close almost nothing. This was a deliberate choice to avoid closing things until the fix was confirmed by the reporter.
But equally that's the first time I've opened the repository in a couple of months and the amount of angst and dread I feel just from the size of that list means I'll probably find yet another excuse not to do anything on it this coming month.
Discussions on this topic feel a lot like "technical solutions to social problems"; by which I mean "well in the ideal world a perfectly logical person would do x, y, z so the system should reflect that". And while a stalebot is the archetypal technical solution to a social problem it at least works with how maintainers work. Sometimes in life you want to ignore a problem and have it go away. When you can't do that, e.g. government bureaucracy, work stuff, social obligations, that's where stress comes from. And asking volunteer maintainers to add a whole new source of stress in their life falls apart when people get busy, or their life circumstances change, or they get ill or tired or whatever.
Yes, in a perfect world the issue backlog would be sacrosanct and perfectly groomed/prioritized. But we're just fleshy sacks of chemicals and we're not perfect. Unrealistic expectations from users are the cause of maintainer burnout.
Because GitHub closed issues are still viewable and searchable (I'd guess most people search it through a search engine not the terrible inbuilt search) I'd disagree that they're deceiving users somehow.
There is framework for everything.
107 projects | /r/ProgrammerHumor | 4 Aug 2022

What about PdfPig? It's under Apache 2.0.
Extract Text from PDF file Blazor
1 project | /r/csharp | 27 Jun 2022

You could try PdfPig. https://uglytoad.github.io/PdfPig/ I've used it for some small tasks and found it very useful. If you want to handle scanned pdfs you would need to use OCR instead.
How to read pdf files in C#?
2 projects | /r/csharp | 2 Dec 2021

PDF Pig is open source and allows you to read text and even extract images.
Add, Remove, Extract and Replace Images in PDF using C#
2 projects | /r/dotnet | 23 Aug 2021

https://uglytoad.github.io/PdfPig/ https://github.com/empira/PDFsharp
Are there any good PDF generation libraries with no paid licensing?
2 projects | /r/dotnet | 15 Mar 2021

Example of document creation API here https://github.com/UglyToad/PdfPig#document-creation-005 and wiki with more details here https://github.com/UglyToad/PdfPig/wiki/Document-Creation
Generating a Report and exporting it as an PDF
1 project | /r/learnprogramming | 12 Mar 2021

Example with PDFpig https://github.com/UglyToad/PdfPig/blob/master/examples/GeneratePdfA2AFile.cs

NetworkX

Posts with mentions or reviews of NetworkX. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2024-03-04.

Routes to LANL from 186 sites on the Internet
1 project | news.ycombinator.com | 4 Mar 2024

Built from this data... https://github.com/networkx/networkx/blob/main/examples/grap...
The Hunt for the Missing Data Type
10 projects | news.ycombinator.com | 4 Mar 2024

I think one of the elements that author is missing here is that graphs are sparse matrices, and thus can be expressed with Linear Algebra. They mention adjacency matrices, but not sparse adjacency matrices, or incidence matrices (which can express muti and hypergraphs).
Linear Algebra is how almost all academic graph theory is expressed, and large chunks of machine learning and AI research are expressed in this language as well. There was recent thread here about PageRank and how it's really an eigenvector problem over a matrix, and the reality is, all graphs are matrices, they're typically sparse ones.
One question you might ask is, why would I do this? Why not just write my graph algorithms as a function that traverses nodes and edges? And one of the big answers is, parallelism. How are you going to do it? Fork a thread at each edge? Use a thread pool? What if you want to do it on CUDA too? Now you have many problems. How do you know how to efficiently schedule work? By treating graph traversal as a matrix multiplication, you just say Ax = b, and let the library figure it out on the specific hardware you want to target.
Here for example is a recent question on the NetworkX repo for how to find the boundary of a triangular mesh, it's one single line of GraphBLAS if you consider the graph as a matrix:
https://github.com/networkx/networkx/discussions/7326
This brings a very powerful language to the table, Linear Algebra. A language spoken by every scientist, engineer, mathematician and researcher on the planet. By treating graphs like matrices graph algorithms become expressible as mathematical formulas. For example, neural networks are graphs of adjacent layers, and the operation used to traverse from layer to layer is matrix multiplication. This generalizes to all matrices.
There is a lot of very new and powerful research and development going on around sparse graphs with linear algebra in the GraphBLAS API standard, and it's best reference implementation, SuiteSparse:GraphBLAS:
https://github.com/DrTimothyAldenDavis/GraphBLAS
SuiteSparse provides a highly optimized, parallel and CPU/GPU supported sparse Matrix Multiplication. This is relevant because traversing graph edges IS matrix multiplication when you realize that graphs are matrices.
Recently NetworkX has grown the ability to have different "graph engine" backends, and one of the first to be developed uses the python-graphblas library that binds to SuiteSparse. I'm not a directly contributor to that particular work but as I understand it there has been great results.
Build the dependency graph of your BigQuery pipelines at no cost: a Python implementation
2 projects | dev.to | 11 Jan 2024

In the project we used Python lib networkx and a DiGraph object (Direct Graph). To detect a table reference in a Query, we use sqlglot, a SQL parser (among other things) that works well with Bigquery.
NetworkX – Network Analysis in Python
1 project | /r/patient_hackernews | 9 Dec 2023

1 project | /r/hackernews | 9 Dec 2023

1 project | /r/hypeurls | 8 Dec 2023

8 projects | news.ycombinator.com | 8 Dec 2023
Custom libraries and utility tools for challenges
1 project | /r/adventofcode | 5 Dec 2023

If you program in Python, can use NetworkX for that. But it's probably a good idea to implement the basic algorithms yourself at least one time.
Google open-sources their graph mining library
7 projects | news.ycombinator.com | 3 Oct 2023

For those wanting to play with graphs and ML I was browsing the arangodb docs recently and I saw that it includes integrations to various graph libraries and machine learning frameworks [1]. I also saw a few jupyter notebooks dealing with machine learning from graphs [2].
Integrations include:
* NetworkX -- https://networkx.org/
* DeepGraphLibrary -- https://www.dgl.ai/
* cuGraph (Rapids.ai Graph) -- https://docs.rapids.ai/api/cugraph/stable/
* PyG (PyTorch Geometric) -- https://pytorch-geometric.readthedocs.io/en/latest/
--
1: https://docs.arangodb.com/3.11/data-science/adapters/
2: https://github.com/arangodb/interactive_tutorials#machine-le...
org-roam-pygraph: Build a graph of your org-roam collection for use in Python
2 projects | /r/orgmode | 7 May 2023

org-roam-ui is a great interactive visualization tool, but its main use is visualization. The hope of this library is that it could be part of a larger graph analysis pipeline. The demo provides an example graph visualization, but what you choose to do with the resulting graph certainly isn't limited to that. See for example networkx.

What are some alternatives?

When comparing PdfPig and NetworkX you can also consider the following projects:

ITextSharp - [DEPRECATED] .NET port of the iText library, only security fixes will be added — please use iText for .NET

Numba - NumPy aware dynamic Python compiler using LLVM

PDFsharp - PDFsharp and MigraDoc Foundation for .NET 6 and .NET Framework

Dask - Parallel computing with task scheduling

Docotic.Pdf - Docotic.Pdf library can create, edit, draw and print PDF files in .NET Core, ASP.NET, Windows Forms, WPF, Xamarin, Blazor, Unity, and HoloLense applications. The library is a 100% managed assembly without unsafe blocks. The assembly has no external dependencies.

julia - The Julia Programming Language

docnet - DocNET is as fast PDF editing and reading library for modern .NET applications

RDKit - The official sources for the RDKit library

Pdfium.Net SDK

snap - Stanford Network Analysis Platform (SNAP) is a general purpose network analysis and graph mining library.

iTextSharp (LGPL / MPL) 4.1.6 for .NET Core - Unofficial .NET Core port of iTextSharp 4.1.6. Last version to be released under the Mozilla Public License and the LGPL.

SymPy - A computer algebra system written in pure Python

PdfPig vs ITextSharp NetworkX vs Numba PdfPig vs PDFsharp NetworkX vs Dask PdfPig vs Docotic.Pdf NetworkX vs julia PdfPig vs docnet NetworkX vs RDKit PdfPig vs Pdfium.Net SDK NetworkX vs snap PdfPig vs iTextSharp (LGPL / MPL) 4.1.6 for .NET Core NetworkX vs SymPy

Compare PdfPig vs NetworkX and see what are their differences.

PdfPig

NetworkX

PdfPig

NetworkX

What are some alternatives?