Framy 1.5 — Video Generation Reaches a New Level
Framy has been updated to version 1.5, and this is no cosmetic patch — it's a major engine upgrade.
What's New
Image Quality
The latent space has been rebuilt with an updated VAE trained on higher-quality data. In practice this means: hair finally looks like hair, text is legible, and fine facial details no longer dissolve into mush. Visual artifacts have been noticeably reduced.
Audio Out of the Box
Framy 1.5 generates synchronized audio in a single pass — dialogue, ambience, sound effects. An updated vocoder has improved speech intelligibility, and cross-modal alignment has reduced lip-sync and timing drift.
Motion and Consistency
The image-to-video mode has been significantly improved: fewer "frozen" frames, fewer spontaneous panning shots, and better preservation of the source image's visual integrity.
Native Vertical Video
For the first time, vertical generation up to 1080×1920 is supported — the model is trained on portrait data rather than cropping from landscape. Reels and Shorts can now be fed directly.
The text encoder has been scaled up 4×, making the model significantly better at understanding complex prompts describing camera angles, character movements, and scene composition.
Prompt Enhancer
Another major bonus is a powerful prompt enhancer. It helps turn a rough request into a more precise and expressive scene description, carefully refining motion, camera angles, dialogue, and atmosphere so the final video lands closer to the original idea on the first try.
In short, Framy 1.5 is one of those cases where the version number is being modest.