python-fake-data-producer-for-apache-kafka
grouparoo
python-fake-data-producer-for-apache-kafka | grouparoo | |
---|---|---|
32 | 27 | |
77 | 607 | |
- | - | |
2.7 | 9.9 | |
7 days ago | about 2 years ago | |
Python | JavaScript | |
Apache License 2.0 | MIT License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
python-fake-data-producer-for-apache-kafka
-
ElephantSQL Is Shutting Down
I had good experience with Aiven in the past, we needed something located in the EU: https://aiven.io/
-
Crossplane: Streamline your infrastructure provisioning & management
Access to Aiven
- Google Cloud Spanner is now half the cost of Amazon DynamoDB
-
Scale up: a MySQL bug story, or why Aiven works
One of the hardest questions we answer for our large enterprise customers is why they should choose Aiven instead of managing their own database and streaming services. It can seem counterintuitive that paying extra for a managed service can save you money. However, when we factor in economies of scale - particularly in regards to access to specialized knowledge and tooling - the case for managed services becomes clear. This was certainly the case for some of our MySQL clients earlier this year, where their investments in Aiven paid off in the form of a quietly managed bug fix.
-
Flink CDC / alternatives
And Kafka + Kafka Connect has https://www.confluent.io/ https://aiven.io/ https://upstash.com/ (and not quite Kafka, but protocol-compatible, https://redpanda.com/)
-
What are your favorite tools or components in the Kafka ecosystem?
Fake data utility - https://github.com/aiven/python-fake-data-producer-for-apache-kafka
-
Do we have such a thing as Postgres Atlas?
For PostgreSQL, similar offerings are: * Google Cloud SQL * AWS RDS * Digital Ocean Postgres * Azure Database for Postgres * Aiven, Instaclustr etc
-
Good database solution
Aiven - https://aiven.io/
-
Why are we paying these folks - a tale of DevRel
Majority of companies layer DevRel on top of marketing as a sort of afterthought, and that’s usually a recipe for failure. All four co-founders at Aiven (the company where I work at) have been long-time open-source maintainers/contributors and highly value the work of DevRel. Similar examples can be seen at HashiCorp, where co-founder Armon Dadgar has been doing DevRel on the whiteboard since the early days of the company. Technical founders know the value of DevRel and know when to form a DevRel team. This is very different from bringing your first DevRel hire onboard and making them convince the leadership why the company needs DevRel in the first place. If your developer advocate needs to explain to the technical leadership the need for DevRel, that's a red flag for that individual and the company.
-
Hetzner continues its growth in the US with a new location
I wonder when Aiven https://aiven.io/ (or something similar) will start supporting hetzner.
grouparoo
-
Reverse ETL recommendations?
Reverse ETL is on AirByte's roadmap under the "Future / Not prioritized" section. I wanted to use Grouparoo as a short term solution, but the repo was archived and I think they stopped taking new cloud customers (unknown if this is wrong/outdated).
-
Reference Data Stack for Data-Driven Startups
There are other tools that we will have to adopt in the future but haven’t yet due to lack of necessity. Specifically, one category that is popular in modern data stacks is Reverse ETL (Hightouch, Census, or Grouparoo). We currently don’t have a usecase for piping data back into 3rd party tools but it will definitely come up in the future.
-
Data pipeline suggestions
Reverse ETL: Grouparoo, Castled
-
Is Reverse ETL a new product or a new ETL/ELT feature?
Grouparoo, the open source Reverse ETL tool we are building, does all of these things. https://www.grouparoo.com
-
Where can I find free data engineering ( big data) projects online?
Ingestion / ETL: Airbyte, Singer, Jitsu Transformation: dbt Orchestration: Airflow, Dagster Testing: GreatExpectations Observability: Monosi Reverse ETL: Grouparoo, Castled Visualization: Lightdash, Superset
- Invite your company
-
Ask HN: Who is hiring? (December 2021)
Grouparoo | Remote (US) | Remote-OK | https://www.grouparoo.com
Grouparoo is a venture-backed software company building open source data tools that make data reliable, accessible, and actionable. We’re empowering teams to make great customer experiences, driven by data. While engineering teams have gotten good at storing and generating data about their customers, it’s rare that this data is used to its full potential in external applications. Grouparoo makes these integrations easy by providing a framework for defining your customer data and reliably syncing it to external tools.
To learn more about who we are, our engineering culture, and whether this is the right place for you, read our Key Values profile: https://www.keyvalues.com/grouparoo
Here are our open roles:
- Senior Backend / Lead Engineer: https://jobs.lever.co/grouparoo/6ba485d1-a5a4-41f0-9fa5-920a...
- Developer Advocate: https://jobs.lever.co/grouparoo/5e1531b4-7ec8-4c10-8e52-fc23...
Tech Stack: TypeScript / Javascript / Node.js, ActionHero, React + Next.js, Postgres & Redis, and whole lot of third-party APIs!
-
Launch HN: Hightouch (YC S19) – Sync data from data warehouses to SaaS tools
Congrats on the launch! Hightouch looks great and this need is real. Things seem to be going well, so I don't think I'm taking too much away by mentioning that we have been been working on Grouparoo, an open source alternative that solves similar pain points.
A few differences: git developer workflow focused (branches, CI, PRs, etc), ability to self host, segmentation in destinations (tagging people in mailchimp based on rules, for example)
https://www.grouparoo.com
-
Reverse ETL
We are building Grouparoo. Obviously, it's a biased sample but we are seeing a few trends at play.
-
What software or coding tools are you trying to get your company to invest in?
Has anyone heard or used of reverse etl or Hightouch or open sourced Grouparoo?
What are some alternatives?
kafka-connect-opensky - Kafka Source Connector reading in from the OpenSky API
rotki - A portfolio tracking, analytics, accounting and management application that protects your privacy
OpenKP - Automatically extracting keyphrases that are salient to the document meanings is an essential step to semantic document understanding. An effective keyphrase extraction (KPE) system can benefit a wide range of natural language processing and information retrieval tasks. Recent neural methods formulate the task as a document-to-keyphrase sequence-to-sequence task. These seq2seq learning models have shown promising results compared to previous KPE systems The recent progress in neural KPE is mostly observed in documents originating from the scientific domain. In real-world scenarios, most potential applications of KPE deal with diverse documents originating from sparse sources. These documents are unlikely to include the structure, prose and be as well written as scientific papers. They often include a much diverse document structure and reside in various domains whose contents target much wider audiences than scientists. To encourage the research community to develop a powerful neural m
airbyte - The leading data integration platform for ETL / ELT data pipelines from APIs, databases & files to data warehouses, data lakes & data lakehouses. Both self-hosted and Cloud-hosted.
fake-data-producer-for-apache-kafka-docker - Fake Data Producer for Aiven for Apache Kafka® in a Docker Image
TileDB - The Universal Storage Engine
Metabase - The simplest, fastest way to get business intelligence and analytics to everyone in your company :yum:
meltano
Grafana - The open and composable observability and data visualization platform. Visualize metrics, logs, and traces from multiple sources like Prometheus, Loki, Elasticsearch, InfluxDB, Postgres and many more.
streamlit - Streamlit — A faster way to build and share data apps.
demo-scene - 👾Scripts and samples to support Confluent Demos and Talks. ⚠️Might be rough around the edges ;-) 👉For automated tutorials and QA'd code, see https://github.com/confluentinc/examples/
PostHog - 🦔 PostHog provides open-source product analytics, session recording, feature flagging and A/B testing that you can self-host.