structured-text-tools
pup
structured-text-tools | pup | |
---|---|---|
13 | 52 | |
6,870 | 8,000 | |
- | - | |
8.1 | 0.0 | |
29 days ago | about 1 month ago | |
HTML | ||
- | MIT License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
structured-text-tools
- Command line tools for manipulating structured text data
-
creating a text file in Linux
This works well in scripts and logs of all the commands you need to do to reproduce the current state of the system from a scratch install. Also can be used with diff -u and patch, sed, perl, and awk oneliners and structured text tools. You can also capture most of the commands using sudo logging feature but it won't capture the here documents. But for modest size files you can use newlines in echo commands. Note that commands which use redrection should use something like ~~~~ sudo bash -c "echo 'foo' >>file.txt" ~~~~ instead of "sudo echo foo >>file.txt" or "echo foo | sudo tee -a file.txt
-
Using Commandline to Process CSV Files
TFA is about how to handle csv files with awk. This might be useful in straightforward cases.
For all others Iād recommend to have a look at
https://github.com/dbohdan/structured-text-tools
which lists tools to handle structure text formats
-
Combine multiple files
in general, I'd pick something from https://github.com/dbohdan/structured-text-tools
- Show HN: Xq ā command-line XML and HTML beautifier and content extractor
- structured-text-tools: A list of command line tools for manipulating structured text data
- A list of command line tools for manipulating structured text data
-
What is your favourite Linux backup software and why?
Also, here is a list of structured text tools. You may find some tools there that are helpful in editing configuration files from the command line. Or you can use "diff -u" to create a patch file (you need to save the patch files along with sudo.log) to recreate. Also, use sfdisk --dump and sfdisk --backup to save partition information in a form that can be used to recreate backups.
pup
-
script to download some notes
And lnk=$(curl -s https://www.selfstudys.com$url |grep "PDFFlip" | cut -d '"' -f 6) to lnk=$(curl -s https://www.selfstudys.com$url | pup "div#PDFF attr{source}" ) here pup will print content of source attribute from div tag with id PDFF i dont know that much about html & css so this is what i came up with. but i am sure you can also select class & make list of suburls from them. check out the video from bugswriter on pup or read docs from git hub for more info github link: https://github.com/ericchiang/pup
-
What monitoring tool do you use or recommend?
jq is pretty amazing. If you are comfortable with its jquery-like CSS selector syntax, then I should also mention a couple similar cli utilities that apply it to HTML: htmlp and pup.
-
Creating a data scraper as a beginner?
Regex is not a great tool for parsing web pages. Open up a browser dev tools window and select a bit of the page. Right click > copy... XPath expression or CSS selector. A proper web scraping tool will accept either of those. No muss, no fuss. You can even use simple command line tools: xpath or pup
- December 5, 2022: FLiP Stack Weekly
-
Show HN: A tool like jq, but for parsing HTML
This is HTML to JSON, written in Rust, and there's also pup[1] which I found out about just the other day on HN[2] which uses a very similar syntax (CSS selectors) but outputs HTML and is written in Go.
I can see room for both though it would interesting to have a more detailed comparison to go on (e.g. types of HTML, speed etc).
[1] https://github.com/ericchiang/pup
[2] https://news.ycombinator.com/item?id=33805732
- Pup: Parsing HTML at the command line
-
pup: Parsing HTML at the Command Line
It looks like the project became inactive for a bit and there are alternatives such as htmlq, etc. https://github.com/ericchiang/pup/issues/150
-
Converting field before delimiter to uppercase and how to replace with multiple newlines
Another tool worth mentioning is pup - it can produce JSON output which means you can pipe it to jq
What are some alternatives?
yq - yq is a portable command-line YAML, JSON, XML, CSV, TOML and properties processor
htmlq - Like jq, but for HTML.
tsv-utils - eBay's TSV Utilities: Command line tools for large, tabular data files. Filtering, statistics, sampling, joins and more.
xidel - Command line tool to download and extract data from HTML/XML pages or JSON-APIs, using CSS, XPath 3.0, XQuery 3.0, JSONiq or pattern matching. It can also create new or transformed XML/HTML/JSON documents.
python-benedict - :blue_book: dict subclass with keylist/keypath support, built-in I/O operations (base64, csv, html, ini, json, pickle, plist, query-string, toml, xls, xml, yaml), s3 support and many utilities.
gron - Make JSON greppable!
concise-encoding - The secure data format for a modern world
yq - Command-line YAML, XML, TOML processor - jq wrapper for YAML/XML/TOML documents
datasette - An open source multi-tool for exploring and publishing data
cascadia - Go cascadia package command line CSS selector
awesome-cli-apps - š„ š š¹ š A curated list of command line apps
ddgr - :duck: DuckDuckGo from the terminal