Cringe

#1
by kabachuha - opened

After more than 1 year since Wan2.1 and practically a year after Wan2.2 you are releasing a specialized model, with nothing novel except "dancing". It's a model on outdated architecture, without inbuilt sound, without any quality improvements.

You are simply mocking the community, which has helped Wan get attention in the first place and to feed upon the dependency.

Ah yes, you can already do audio-video sync in LTX, simply providing the latent audio channel. Up to ~30 seconds, on consumer hardware. I don't think your project here is something special.

LTX, Nvidia and Sber's Kandinsky are building audio-video models, the latter dropping at any moment now. This is the last chance to save your face.

Alibaba has already lost the local image market. It is only a matter of time before we permanently replace Wan.

This is the last chance to save your face.

u sure 100% positive u r not on drugs?

Want a Wan that can generate audio? Well, it already exists Mova-720p / Mova-360p. It's based on Wan2.2, and the audio part uses HunyuanVideo-Foley.

Please remove this cringe post 😟

slop ai

yeah well, if you get something for free and dont like it you could just shut up..

After more than 1 year since Wan2.1 and practically a year after Wan2.2 you are releasing a specialized model,

People with resources do research and post models and code to push the science. It's not to cater to ungrateful fucks like yourself.

You're not special - go back to reddit.

Alibaba has already lost the local image market. It is only a matter of time before we permanently replace Wan.

You think they care? And do you think if people who work w/ development/scholar teams training models that take massive GPU clusters are envious of the type of nonsense you're posting here?

No they aren't

It's always the stupidest who are the loudest - I'm pushing 50yrs old, maybe one day you'll grow the fuck up kid.

buddy. stop expecting companies to just do whatever you want. if you could do it yourself you wouldnt be complaining. so let the big guys do their thing. go complain on twitter and let it out lol respectfully

I don't think your project here is something special.
with nothing novel except "dancing"

See... I do think Alibaba has better things to do with their compute besides whatever this is. (unless this is video gen research, which if it is, then that's a useful thing to do with compute)
But I do think it's special, despite me reacting so negatively to this model as well. I don't think Alibaba cares about who kabachuha is.
If this is for the sake of ML advancements, then as much as I do hate a model like this, I understand this is research material.
So... it probably is something special!

It's a model on outdated architecture, without inbuilt sound, without any quality improvements.

Dude. You're literally getting the weights for FREE. There's always the option to just... never use the weights and move on with your life.

You are simply mocking the community,

Please never speak on behalf for us ever again. No one is being mocked, you're mocking yourself. My younger cousin in ELEMENTARY SCHOOL doesn't throw tantrums like this.

This is the last chance to save your face.

Christ. Who is bro?!?!?!😭
Go complain on Reddit instead bro...

I mean... personally, I'm willing to forgive them for this horrible sin if they release WanStreamer V1 + V2 + V3 very soon :)

Wan's policy on open-source development is completely different from what you might think. Wan 2.1 and 2.2 were released primarily as a foundation for scientific research and the development of academic papers. These papers develop the Wan ecosystem itself for free, laying the groundwork for Alibaba's own labs to build closed-source commercial applications and generate profit.

Why are we getting Wan-Dancer instead of Wan 2.7? Because in 99% of cases, the Wan 2.1 or Wan 2.2 architecture is already chosen as the basis for conducting scientific research anyway.

Grumbling at the creators of Wan for releasing what we don't want or not releasing what we do want is pointless, because Alibaba doesn't care anyway. We will just keep generating dancing waifus on Wan 2.1 or Wan 2.2, just as we always have.

What could break this vicious cycle? Here is my take:

  1. The release of a radically better open-source model by a company other than Alibaba. And this is where you might think of LTX 2.3. I think everyone would agree that LTX is worse than Wan; it has a lot of drawbacks. LTX is only suitable for generating hyperrealism and talking heads, and it requires mile-long prompts just to get a character to even raise an eyebrow. Additionally, LTX has the "LTX-2 Community License Agreement," which is not very conducive to the development of academic papers around this model. From what I can see, LTX does not plan to release new models, choosing instead to build their ecosystem increasingly around LTX 2.3. However, perhaps a new player besides Wan and LTX will emerge to shift the balance of power between them, thereby forcing Wan to reclaim its leadership by releasing new models.

  2. Alibaba's cutting-edge developments will diverge too far from the Wan 2.2 and 2.1 architectures. This would make the application of academic papers written for Wan 2.1 and Wan 2.2 impossible. Such a gap would force Alibaba to release an updated Wan X.X architecture so they can continue to reap the fruits of independent enthusiast scientists' labor.

I don't see any other scenarios in which Alibaba and the Wan team would decide to release anything better than Wan 2.2 or Wan 2.1.

I hope you find my arguments convincing. I don't think they are pessimistic or optimistic. Like you, I want to see new open-source SOTA models, but I am also trying to analyze trends critically and unemotionally.

If you have any counterarguments, please feel free to share them.

Wise words from a decent man right there

  1. The release of a radically better open-source model by a company other than Alibaba. And this is where you might think of LTX 2.3. I think everyone would agree that LTX is worse than Wan; it has a lot of drawbacks. LTX is only suitable for generating hyperrealism and talking heads, and it requires mile-long prompts just to get a character to even raise an eyebrow. Additionally, LTX has the "LTX-2 Community License Agreement," which is not very conducive to the development of academic papers around this model. From what I can see, LTX does not plan to release new models, choosing instead to build their ecosystem increasingly around LTX 2.3. However, perhaps a new player besides Wan and LTX will emerge to shift the balance of power between them, thereby forcing Wan to reclaim its leadership by releasing new models.

Seconded. Wan2.2 is a better model than LTX2.3 in terms of flexibility towards training, learning new concepts, prompt adherence and understanding. If we speak in terms of tokens, LTX2.3 will end up using a lot more than Wan2.2 in retries just to get the desired output. LTX2.3 has its strengths in hyperrealism and audio generation, but from a pure research perspective the malleability of Wan2.2 towards learning is amazing.

这不是产品,这是个开源项目。即使有一天他们不再开源任何东西了,我们也只能表示失望,而没有理由去谴责他们

Sign up or log in to comment