DALLE2-pytorch
VQGAN-CLIP
DALLE2-pytorch | VQGAN-CLIP | |
---|---|---|
65 | 67 | |
10,826 | 2,563 | |
- | - | |
6.8 | 0.0 | |
3 months ago | over 1 year ago | |
Python | Python | |
MIT License | GNU General Public License v3.0 or later |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
DALLE2-pytorch
-
One year ago I got access to closed beta DALL-E 2.
I was showing people Dalle2 last year and telling them how much of an impact an open source solution was going to have on, well, everything to do with art and design. (At the time Stable Diffusion had not released, not even the leak, and all hopes was on https://github.com/lucidrains/DALLE2-pytorch)
- [Machinelearning] [D] Quelqu'un travaille-t-il sur l'open-sourcing de Dall-E 2 ?
-
AMA (Emad here hello)
Stable diffusion is the model, MJ will use a variant and DALL-E is the old version (we have our own implementation from our distinguished fellow Lucidrains here: https://github.com/lucidrains/DALLE2-pytorch)
-
An impressionist painting of an floating raccoon god, 4k, digital painting, trending on artstation
Sadly I don't think so. From what I understand the architecture is fixed to 1024x1024 pictures.
- I asked AI to turn P&R characters into muppets..
-
Comparison of AI text-to-image generators
The code is open source, the model is not I believe. https://github.com/lucidrains/DALLE2-pytorch
- Protests erupt outside of DALL-E offices after pricing implementation, press photograph
-
$15 for 115 “generation increments” Very expensive Beta pricing announcement. Dissapointed
Phil Wang has been fairly prolific at creating open source implementations of these text to image models. For example, here is the dalle-2 repo https://github.com/lucidrains/DALLE2-pytorch
-
DALL·E Now Available in Beta
There's already an open-source implementation of DALL-E 2 (https://github.com/lucidrains/DALLE2-pytorch) and a pretrained model for it should be released within this year.
Also true for Google's Imagen, which should be even better than DALLE-2 (and faster) https://github.com/lucidrains/imagen-pytorch.
This is possible because the original research papers behind both DALLE-2 and Imagen were publicly released.
-
would love to know what portion of this prompt is not allowed
The paper describing the model is public and has been implemented here, but that's not the hard part. The model likely requires months of compute and dozens of gigabytes of VRAM to train and run, likely costing several hundred thousand dollars.
VQGAN-CLIP
-
📚 Tutorials & 🎨 AI Art Generation Tool List Mega Thread
VQGAN-CLIP
-
Which is your favorite text to image model overall?
I've screwed with many text-to-image models over the past couple of years, and I found that while I currently enjoy Stable Diffusion's coherency, I have a soft spot for the ImageNet model used by default for VQGAN+CLIP. It easily approaches the uncanny valley when generating people or animals, but makes for great abstract backgrounds and wallpapers. I already have nostalgia for generating images with it on my CPU overnight.
-
Stable Diffusion Announcement
For someone only tangentially familiar with this space, how is this different than e.g. https://github.com/nerdyrodent/VQGAN-CLIP which you can also run at home? Is it the quality of the generated images?
-
Medieval Noir - VQGAN-CLIP - COCO Checkpoint
Used https://github.com/nerdyrodent/VQGAN-CLIP
- Once have access, do you run it on your computer or over the internet on Open-AI's computers?
- How to get AI imaging effect in Premiere pro
-
A Guide to Asking Robots to Design Stained Glass Windows
I don't have any of the DALL-Es but I do have a couple from github [1], [2] which gave these outputs[3]
[1] https://github.com/nerdyrodent/VQGAN-CLIP
-
How not to waste $1600?
If you want to try your hand at buggering your whole system - try playing with AI image generation as it uses all possible computer assets :D . There is a lot of forms and installations for those but I VQGANs from github the easiest. Problem is that some require familarity with shell, python and in some cases - you need to enable the Linux subsystem in Windows (is it called a subsystem? it is not exactly a VM). This one is the easiest to install out of all I tried. But I liked the results of Pixray most but I wrecked it. I use this one nowadays.
- Ask HN: Is there a publicly available (not private beta) text-to-image API?
-
Got a Machine Learning Algorithm to depict Aphex
For those that are interested, I used VQGAN-CLIP, specifically this GitHub repository
What are some alternatives?
dalle-mini - DALL·E Mini - Generate images from a text prompt
CLIP-Guided-Diffusion - Just playing with getting CLIP Guided Diffusion running locally, rather than having to use colab.
disco-diffusion
DALLE-mtf - Open-AI's DALL-E for large scale training in mesh-tensorflow.
DALLE-pytorch - Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
image-super-resolution - 🔎 Super-scale your images and run experiments with Residual Dense and Adversarial Networks.
DALL-E - PyTorch package for the discrete VAE used for DALL·E.
deep-daze - Simple command line tool for text to image generation using OpenAI's CLIP and Siren (Implicit neural representation network). Technique was originally created by https://twitter.com/advadnoun
dalle-2-preview
waifu2x - Image Super-Resolution for Anime-Style Art
latent-diffusion - High-Resolution Image Synthesis with Latent Diffusion Models
stable-diffusion - A latent text-to-image diffusion model