I am new to this interesting place. I just wanted to give a brief rundown of how I got here and my experience with AI so far. I first started out using Flux Dev through Stability Matrix. I was happy enough for a while using their general inference panel and adding a few LoRA here and there. Then one day I made a shocking discovery; the Open Web UI button. It turned out, much to my surprise, that I had essentially been using a middle man the entire time. My head swam trying to figure out what was going on in the ComfyUI interface.
What ultimately forced me to get to know ComfyUI was laziness. I noticed in ComfyUI that there was a Power LoRA loader node. It allowed me to do text search of all the LoRA on my computer and open an info panel that pulled trigger words from this site. In Stability Matrix it was just a giant drop down menu with no info access. At this point I had already accumulated 100ish LoRA and could no longer stand to deal with the pain. Thus began my long journey into the deep end. I started with a generic workflow and soon I was happily pushing the limits of every possible setting and breaking things completely. I couldn't have been happier.
Over the next two months during my summer vacation from college I spent roughly 8 hours a day working with Flux, Wan 2.1 and 2.2 and eventually Pony. I love Flux for its natural language understanding and realism. I can feed Flux a book and it gives me an unreasonably accurate image given the quantity of madness I can push into its model. Additionally, Flux's ability to process complex prompts involving multiple people and intricate location dynamics is miles ahead of everything else. I hate Flux for its rigidity and requiring a LoRA for basically everything. I will never understand how a model can both be so intensely powerful and dynamic and also be so suffocatingly rigid.
Wan is, well, Wan. Wan 2.1 was a great model to learn and grow my first videos and Wan 2.2 is developing into a great video model. As expected there have been some growing pains and I am so incredibly grateful to the folks doing the hard work of fine-tuning the LoRA development process. Some truly incredible things are being released daily. As much as I love Wan 2.2 I am still finding myself wishing it had CLIP Vision support but who knows what the future will hold.
Pony is something I had never planned on using but now I use it almost exclusively. I love Pony for it's ability to produce everything from an uncanny beautiful hyperrealism to anime and beyond. I tend to stick almost exclusively to the realism side of things and this was one of the reasons why I never planned on using Pony. However, I kept seeing these videos and images that had this intense gritty raw realism. A kind of hyperrealism that seemed impossible to capture even in real life. Every one of them turned out to be made not in Flux, but in Pony. I had to set aside my preconceived biases about Pony and see what the model could do and it was a game changer.
However... I hate Pony's prompting with a passion. After using Flux for roughly 200 hours trying to work with Pony was a nightmare. I have gotten a bit more used to how it works but still routinely turn to a guide or to images generated with prompts provided as templates. I know I will eventually get used to it but for now... all we know is pain. In theory Pony's prompting is much more simple however it is incredibly limiting coming from a natural language model. Its ability to make images with more than one person with unique characteristics or heaven forbid an interesting location beyond generic one word tags is abysmal and infuriating. But I just can't give you up, Pony.
This brings me to my closing thoughts, and anyone that has made it this far is a legend. A huge thank you to everyone that actually provides accurate prompts with their work and cooks in the metadata. I have learned so much interesting info from peoples workflows and prompts. Now that I am submitting my own work for public view on this site I am returning the favor and hopefully some newbies can find some interesting tidbits from my stuff.
The only reason I would not want to share everything about how the sausage is made is simply because I would prefer not to give anyone an easy button. The best times and biggest growth I have had in image and video generation have been from completely breaking my workflow or doing something totally different that I hadn't seen anyone else try and learning from that. Providing an already fully setup and tweaked workflow that only requires a prompt and a button press robs people of that learning process.
Big takeaways:
Experiment ALOT, break everything, try everything.
If you are serious about this stuff make it a priority to generate on your own hardware.
Archive everything. This space is tightening constantly, you never know when a model or LoRA will vanish overnight. I have 700 LoRA on local storage.
Try a wide variety of Sampler/Scheduler combos, they can completely change the dynamics and outcome of a prompt. Living in the walled garden of Euler A and Beta is a death sentence to style stagnation.
Learn from others work and workflows but make your own from what you learn.
Don't get scared away by the VRAM trap. Since I knew nothing about limitations or the widespread dogma in the community I was using full fp32 models in Flux and standard Wan 2.1 models on a card with 12GB of VRAM and 32GB of system ram. Some make things like this out to be impossible. In my experience system ram was a bigger bottleneck. I have upgraded to 80GB of system ram and now that the models can happily load in and out of ram I can generate video and images to my hearts content while multitasking on my PC. You don't need a 5090.
PC specs: 11th gen i9, 80GB ddr4 RAM, 4070ti 12GB

