Semantic-search-tweets Alternatives

Similar projects and alternatives to semantic-search-tweets

marqo

114 4,124 9.3 Python semantic-search-tweets VS marqo

Unified embedding generation and search engine. Also available on cloud - cloud.marqo.ai
gpt4-pdf-chatbot-langchain

32 14,573 3.9 TypeScript semantic-search-tweets VS gpt4-pdf-chatbot-langchain

GPT4 & LangChain Chatbot for large PDF docs
InfluxDB

www.influxdata.com featured

Power Real-Time Data Analytics at Scale. Get real-time insights from all types of time series data with InfluxDB. Ingest, query, and analyze billions of data points in real-time with unbounded cardinality.
Fast_Sentence_Embeddings

3 603 0.0 Jupyter Notebook semantic-search-tweets VS Fast_Sentence_Embeddings

Compute Sentence Embeddings Fast!

NOTE: The number of mentions on this list indicates mentions on common posts plus user suggested alternatives. Hence, a higher number means a better semantic-search-tweets alternative or higher similarity.

Suggest an alternative to semantic-search-tweets

semantic-search-tweets reviews and mentions

Posts with mentions or reviews of semantic-search-tweets. We have used some of these posts to build our list of alternatives and similar projects. The last one was on 2023-03-30.

You probably shouldn't use OpenAI's embeddings
5 projects | news.ycombinator.com | 30 Mar 2023

It's in the repo:
You first create embeddings. What is this? It's an n-dimensional vector space with your tweets 'embedded' in that space. Each word is an n-dimensional vector in this space. The vectorization is supposed to maintain 'semantic distance'. Basically, if two words are very close in meaning or related (by say frequently appearing next to each other in corpus) they should be 'close' in some of those n-dimensions as well. The result at the end is the '.bin' file, the 'semantic model' of your corpus.
https://github.com/dbasch/semantic-search-tweets/blob/main/e...
For semantic search, you run the same embedding algorithm against the query and take the resultant vectors and do similarity search via matrix ops, resulting in a set of results, with probabilities. These point back to the original source, here the tweets, and you just print the tweet(s) that you select from that result set.
https://github.com/dbasch/semantic-search-tweets/blob/main/s...
Experts can chime in here but there are knobs such as 'batch size' and the functions you use to index. (cosine was used here.)
So the various performance dimensions of the process should also be clear. There is a fixed cost of making the embeddings of your data. There is a per-op embedding of your query, and then running the similarity algorithm to find the result set.