• Nosotros
  • Publicidad
  • Trabaja con nosotros
  • Contactos
miércoles, agosto 19, 2026
  • Login
No Result
View All Result
NEWSLETTER
Despertar Matinal
  • Titulares del Día
    • All
    • En Portada
    Teobaldo Durán: “MP ocultó responsables explosión en SC; desconoce informe J-2; No descarta nuevas acciones”

    Teobaldo Durán: “MP ocultó responsables explosión en SC; desconoce informe J-2; No descarta nuevas acciones”

    Movimientos Médicos denuncian crisis del sector salud y llaman a rechazar complicidad entre autoridades del CMD y el Gobierno

    Movimientos Médicos denuncian crisis del sector salud y llaman a rechazar complicidad entre autoridades del CMD y el Gobierno

    Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

    Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

    Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

    Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

    Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

    Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

    Ito Bisonó asume como ministro de Relaciones Exteriores con una trayectoria de gestión pública y amplios vínculos internacionales

    Ito Bisonó asume como ministro de Relaciones Exteriores con una trayectoria de gestión pública y amplios vínculos internacionales

    Raúl Martínez: seis años bastan para exigir resultados

    Raúl Martínez: seis años bastan para exigir resultados

    Procurador fiscal pide aumento salarial para representantes del Ministerio Público

    Procurador fiscal pide aumento salarial para representantes del Ministerio Público

    El Instituto Duartiano aboga por preservar la autodeterminación de RD ante versiones sobre presiones de EE. UU.

    El Instituto Duartiano aboga por preservar la autodeterminación de RD ante versiones sobre presiones de EE. UU.

    Trending Tags

    • Mundo
      • All
      • América Latina
      • Conflictos Internacionales
      • Estados Unidos
      • Europa
      • Geopolítica
      • Haití
      • Medio Oriente
      Clarín publicó una noticia del cierre de una cadena de supermercados sin mencionar que es de EEUU

      Clarín publicó una noticia del cierre de una cadena de supermercados sin mencionar que es de EEUU

      Boca encamina dos nuevas salidas: Marcelo Weigandt y Juan Barinaga se irían a préstamo

      Boca encamina dos nuevas salidas: Marcelo Weigandt y Juan Barinaga se irían a préstamo

      GTA 6: así sería el mapa filtrado de Vice City y todo Leonida

      GTA 6: así sería el mapa filtrado de Vice City y todo Leonida

      La UCR de Córdoba va a internas por desacuerdos en las bancas del Congreso partidario

      La UCR de Córdoba va a internas por desacuerdos en las bancas del Congreso partidario

      Moderna dispara sus acciones más del 100% tras el éxito de su vacuna contra el cáncer

      Moderna dispara sus acciones más del 100% tras el éxito de su vacuna contra el cáncer

      La Unión Europea instó a que todos los inmigrantes ilegales en Ceuta sean devueltos a Marruecos

      La Unión Europea instó a que todos los inmigrantes ilegales en Ceuta sean devueltos a Marruecos

      Qué se sabe sobre el futuro futbolístico de Mauro Icardi: los clubes que estarían dispuestos a aceptarlo

      Qué se sabe sobre el futuro futbolístico de Mauro Icardi: los clubes que estarían dispuestos a aceptarlo

      Tesla se prepara para lanzar el Cybercab, pero crecen las dudas sobre si está listo

      Tesla se prepara para lanzar el Cybercab, pero crecen las dudas sobre si está listo

      La inversión de Peter Thiel en Vaca Muerta es la segunda mas grande de su cartera

      La inversión de Peter Thiel en Vaca Muerta es la segunda mas grande de su cartera

      Trending Tags

      • Nacionales
        • All
        • Bávaro Punta Cana
        • Educación
        • Gobierno
        • Infraestructura
        • Justicia
        • Obras Públicas
        • Opinión
        • Provincias
        • Seguridad Ciudadana
        • semana santa 2026
        • Sociedad
        • Transporte
        Indotel adjudica a Viettel y Claro Dominicana bloques de frecuencias del espectro radioeléctrico

        Indotel adjudica a Viettel y Claro Dominicana bloques de frecuencias del espectro radioeléctrico

        Teobaldo Durán: “MP ocultó responsables explosión en SC; desconoce informe J-2; No descarta nuevas acciones”

        Teobaldo Durán: “MP ocultó responsables explosión en SC; desconoce informe J-2; No descarta nuevas acciones”

        Juan Manuel Méndez García asume dirección del Intrant

        Juan Manuel Méndez García asume dirección del Intrant

        Movimientos Médicos denuncian crisis del sector salud y llaman a rechazar complicidad entre autoridades del CMD y el Gobierno

        Movimientos Médicos denuncian crisis del sector salud y llaman a rechazar complicidad entre autoridades del CMD y el Gobierno

        Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

        Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

        Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

        Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

        Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

        Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

        CRR Las Parras crea talleres industriales de producción de colchones, ropa y tapicería

        CRR Las Parras crea talleres industriales de producción de colchones, ropa y tapicería

        Alejandro Campos es juramentado por Eduardo Estrella como...

        Alejandro Campos es juramentado por Eduardo Estrella como…

        Trending Tags

        • Política
          • All
          • Congreso
          • Opinión Política
          • Partidos Políticos
          • Poder Municipal
          • Transparencia y Corrupción
          FP pide interpelar al ministro de Educación Luis Miguel De Camps por dificultades previo al inicio del año escolar

          FP pide interpelar al ministro de Educación Luis Miguel De Camps por dificultades previo al inicio del año escolar

          Colombia Alcántara será moderadora del XIX congreso...

          Colombia Alcántara será moderadora del XIX congreso…

          Milton Morrison reafirma alianza con Abinader y anuncia nueva...

          Milton Morrison reafirma alianza con Abinader y anuncia nueva…

          Empresarios de Hato Mayor expresan respaldo a Leonel Fernández y fortalecen proyecto político rumbo a 2028

          Empresarios de Hato Mayor expresan respaldo a Leonel Fernández y fortalecen proyecto político rumbo a 2028

          PRM en Santo Domingo Norte resalta gestión del presidente...

          PRM en Santo Domingo Norte resalta gestión del presidente…

          ARTICULO: De los millones de seguidores al poder: gobernar un país no es hacer un reality en YouTube

          ARTICULO: De los millones de seguidores al poder: gobernar un país no es hacer un reality en YouTube

          Estados Unidos no descarta operación militar contra Cuba

          Estados Unidos no descarta operación militar contra Cuba

          Tribunal Constitucional ratifica que País Posible es la 7ma fuerza...

          Tribunal Constitucional ratifica que País Posible es la 7ma fuerza…

          Sismo en Colombia suma 181 fallecidos

          Sismo en Colombia suma 181 fallecidos

          Trending Tags

          • Deportes
            • All
            • Atletas Dominicanos
            • Béisbol
            DR Open Kiteboarding Championship reúne atletas de 15 países y reafirma a Cabarete como capital del kitesurf del Caribe

            Cabarete se corona como capital histórica del kitesurf con el DR Open Championship 2026

            El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

            El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

            La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

            La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

            Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

            Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

            Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

            Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

            Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

            Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

            Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

            Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

            30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

            Buffalo recibe a Montreal para abrir la segunda ronda

            Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

            Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

            Trending Tags

            • Economía
              • All
              • Combustibles
              • Energía
              • Indicadores Económicos
              • Sector Energético
              • Turismo
              Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

              Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

              Aventúrate RD 2026

              Aventúrate RD 2026 revela agenda oficial y consolida el turismo de aventura dominicano

              WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

              WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

              Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

              Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

              Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

              Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

              Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

              Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

              El ministerio de Turismo cierra con éxito festival gastronómico “Saborea el Paraíso” en Sánchez, Samaná

              El Ministerio de Turismo celebra un exitoso cierre del festival gastronómico «Saborea el Paraíso» en Sánchez, Samaná

              El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

              El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

              Más allá del comercio: los efectos del arancel estadounidense sobre el turismo dominicano

              Arancel de EE.UU. pone a prueba al turismo dominicano y al silencio oficial del gobierno

              Trending Tags

              • Ciencia
                • All
                • Energía
                • Innovación
                • Investigación Científica
                • Salud y Medicina
                • Tecnología Médica
                81-year-old admits German cold-case murder of US tourist in 1994

                81-year-old admits German cold-case murder of US tourist in 1994

                FP podría ganar elecciones en primera vuelta, según Bauta Rojas

                FP podría ganar elecciones en primera vuelta, según Bauta Rojas

                Lindsay Clancy's trial put postpartrum psychosis on the map. Experts want progress to follow

                Lindsay Clancy’s trial put postpartrum psychosis on the map. Experts want progress to follow

                Israel opens up tenders for controversial West Bank settlement project

                Israel opens up tenders for controversial West Bank settlement project

                Target receives $1bn boost from Trump tariff refunds

                Target receives $1bn boost from Trump tariff refunds

                Ukrainian man arrested in Croatia over Nord Stream blasts

                Ukrainian man arrested in Croatia over Nord Stream blasts

                Helicopter crashes at Mount Ololokwe in Kenya, killing six people

                Helicopter crashes at Mount Ololokwe in Kenya, killing six people

                Apparent human remains found in US reservoir as water levels hit record low

                Apparent human remains found in US reservoir as water levels hit record low

                Erin Patterson appeal: Mushroom murderer's trial undermined by hotel mix-up, court told

                Erin Patterson appeal: Mushroom murderer’s trial undermined by hotel mix-up, court told

                Trending Tags

                • Tecnología
                  • All
                  • Aplicaciones
                  • Inteligencia Artificial
                  ¿Qué es mejor que ver eclipses desde el espacio? Prueba las luces del norte y del sur.

                  ¿Qué es mejor que ver eclipses desde el espacio? Prueba las luces del norte y del sur.

                  30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                  Un telescopio espía la estrella más rápida de la Vía Láctea girando alrededor de un agujero negro

                  VentureBeat nombra a Rob Strechay como su primer analista principal, ampliando su impulso de investigación de IA empresarial

                  VentureBeat nombra a Rob Strechay como primer analista principal, ampliando el esfuerzo de investigación de IA empresarial

                  VentureBeat nombra a Rob Strechay como su primer analista principal, ampliando su impulso de investigación de IA empresarial

                  VentureBeat nombra a Rob Strechay como su primer analista principal, ampliando su impulso de investigación de IA empresarial

                  Desde perros robot hasta ayudantes, China muestra sus ambiciones en materia de robótica en una conferencia

                  Desde perros robot hasta ayudantes, China muestra sus ambiciones en materia de robótica en una conferencia

                  30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                  Fascinantes medusas ocupan un lugar central en el museo de EE. UU. dedicado a estas criaturas

                  China recupera la etapa de un cohete en tierra por primera vez, después de una recuperación anterior en el mar

                  China recupera la etapa de un cohete en tierra por primera vez, después de una recuperación anterior en el mar

                  Un almacén arde cerca de Moscú. En Ucrania, los equipos de drones se preparan para atacar nuevamente

                  Un almacén arde cerca de Moscú. En Ucrania, los equipos de drones se preparan para atacar nuevamente

                  Las acciones del fabricante chino de robots humanoides Unitree se disparan en su debut comercial en Shanghai

                  Las acciones del fabricante chino de robots humanoides Unitree se disparan en su debut comercial en Shanghai

                  Trending Tags

                  • Entretenimiento
                    • All
                    • Cine y Series
                    • Cultura Digital
                    • Cultura Popular
                    • Gastronomía
                    • Música
                    Después de perder a un amigo y escribir 'Say So', Dan + Shay regresan con la autobiográfica 'Young'

                    Después de perder a un amigo y escribir ‘Say So’, Dan + Shay regresan con la autobiográfica ‘Young’

                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                    El Centro Kennedy dice a la corte que no intentará restaurar el nombre de Trump en el edificio antes del 8 de septiembre

                    Patólogo forense detalla las heridas fatales de Tupac Shakur en el juicio de Duane 'Keffe D' Davis

                    Patólogo forense detalla las heridas fatales de Tupac Shakur en el juicio de Duane ‘Keffe D’ Davis

                    El cofundador de ESPN, Bill Rasmussen, muere a los 93 años por los efectos de la enfermedad de Parkinson

                    El cofundador de ESPN, Bill Rasmussen, muere a los 93 años por los efectos de la enfermedad de Parkinson

                    Fox Sports transmitirá 35 partidos de voleibol femenino, incluidos 8 en Fox

                    Fox Sports transmitirá 35 partidos de voleibol femenino, incluidos 8 en Fox

                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                    Una película de animación china calificada de «terrible» se convierte en un éxito de taquilla

                    Shakira realiza visita sorpresa a Colombia afectada por el terremoto y se compromete a construir nuevas escuelas

                    Shakira realiza visita sorpresa a Colombia afectada por el terremoto y se compromete a construir nuevas escuelas

                    Bonnie Tyler es recordada como estrella mundial en su funeral en Gales

                    Bonnie Tyler es recordada como estrella mundial en su funeral en Gales

                    Dave Marsh, biógrafo y crítico musical de Bruce Springsteen, muere a los 76 años

                    Dave Marsh, biógrafo y crítico musical de Bruce Springsteen, muere a los 76 años

                    Trending Tags

                    • Titulares del Día
                      • All
                      • En Portada
                      Teobaldo Durán: “MP ocultó responsables explosión en SC; desconoce informe J-2; No descarta nuevas acciones”

                      Teobaldo Durán: “MP ocultó responsables explosión en SC; desconoce informe J-2; No descarta nuevas acciones”

                      Movimientos Médicos denuncian crisis del sector salud y llaman a rechazar complicidad entre autoridades del CMD y el Gobierno

                      Movimientos Médicos denuncian crisis del sector salud y llaman a rechazar complicidad entre autoridades del CMD y el Gobierno

                      Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

                      Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

                      Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

                      Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

                      Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

                      Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

                      Ito Bisonó asume como ministro de Relaciones Exteriores con una trayectoria de gestión pública y amplios vínculos internacionales

                      Ito Bisonó asume como ministro de Relaciones Exteriores con una trayectoria de gestión pública y amplios vínculos internacionales

                      Raúl Martínez: seis años bastan para exigir resultados

                      Raúl Martínez: seis años bastan para exigir resultados

                      Procurador fiscal pide aumento salarial para representantes del Ministerio Público

                      Procurador fiscal pide aumento salarial para representantes del Ministerio Público

                      El Instituto Duartiano aboga por preservar la autodeterminación de RD ante versiones sobre presiones de EE. UU.

                      El Instituto Duartiano aboga por preservar la autodeterminación de RD ante versiones sobre presiones de EE. UU.

                      Trending Tags

                      • Mundo
                        • All
                        • América Latina
                        • Conflictos Internacionales
                        • Estados Unidos
                        • Europa
                        • Geopolítica
                        • Haití
                        • Medio Oriente
                        Clarín publicó una noticia del cierre de una cadena de supermercados sin mencionar que es de EEUU

                        Clarín publicó una noticia del cierre de una cadena de supermercados sin mencionar que es de EEUU

                        Boca encamina dos nuevas salidas: Marcelo Weigandt y Juan Barinaga se irían a préstamo

                        Boca encamina dos nuevas salidas: Marcelo Weigandt y Juan Barinaga se irían a préstamo

                        GTA 6: así sería el mapa filtrado de Vice City y todo Leonida

                        GTA 6: así sería el mapa filtrado de Vice City y todo Leonida

                        La UCR de Córdoba va a internas por desacuerdos en las bancas del Congreso partidario

                        La UCR de Córdoba va a internas por desacuerdos en las bancas del Congreso partidario

                        Moderna dispara sus acciones más del 100% tras el éxito de su vacuna contra el cáncer

                        Moderna dispara sus acciones más del 100% tras el éxito de su vacuna contra el cáncer

                        La Unión Europea instó a que todos los inmigrantes ilegales en Ceuta sean devueltos a Marruecos

                        La Unión Europea instó a que todos los inmigrantes ilegales en Ceuta sean devueltos a Marruecos

                        Qué se sabe sobre el futuro futbolístico de Mauro Icardi: los clubes que estarían dispuestos a aceptarlo

                        Qué se sabe sobre el futuro futbolístico de Mauro Icardi: los clubes que estarían dispuestos a aceptarlo

                        Tesla se prepara para lanzar el Cybercab, pero crecen las dudas sobre si está listo

                        Tesla se prepara para lanzar el Cybercab, pero crecen las dudas sobre si está listo

                        La inversión de Peter Thiel en Vaca Muerta es la segunda mas grande de su cartera

                        La inversión de Peter Thiel en Vaca Muerta es la segunda mas grande de su cartera

                        Trending Tags

                        • Nacionales
                          • All
                          • Bávaro Punta Cana
                          • Educación
                          • Gobierno
                          • Infraestructura
                          • Justicia
                          • Obras Públicas
                          • Opinión
                          • Provincias
                          • Seguridad Ciudadana
                          • semana santa 2026
                          • Sociedad
                          • Transporte
                          Indotel adjudica a Viettel y Claro Dominicana bloques de frecuencias del espectro radioeléctrico

                          Indotel adjudica a Viettel y Claro Dominicana bloques de frecuencias del espectro radioeléctrico

                          Teobaldo Durán: “MP ocultó responsables explosión en SC; desconoce informe J-2; No descarta nuevas acciones”

                          Teobaldo Durán: “MP ocultó responsables explosión en SC; desconoce informe J-2; No descarta nuevas acciones”

                          Juan Manuel Méndez García asume dirección del Intrant

                          Juan Manuel Méndez García asume dirección del Intrant

                          Movimientos Médicos denuncian crisis del sector salud y llaman a rechazar complicidad entre autoridades del CMD y el Gobierno

                          Movimientos Médicos denuncian crisis del sector salud y llaman a rechazar complicidad entre autoridades del CMD y el Gobierno

                          Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

                          Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

                          Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

                          Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

                          Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

                          Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

                          CRR Las Parras crea talleres industriales de producción de colchones, ropa y tapicería

                          CRR Las Parras crea talleres industriales de producción de colchones, ropa y tapicería

                          Alejandro Campos es juramentado por Eduardo Estrella como...

                          Alejandro Campos es juramentado por Eduardo Estrella como…

                          Trending Tags

                          • Política
                            • All
                            • Congreso
                            • Opinión Política
                            • Partidos Políticos
                            • Poder Municipal
                            • Transparencia y Corrupción
                            FP pide interpelar al ministro de Educación Luis Miguel De Camps por dificultades previo al inicio del año escolar

                            FP pide interpelar al ministro de Educación Luis Miguel De Camps por dificultades previo al inicio del año escolar

                            Colombia Alcántara será moderadora del XIX congreso...

                            Colombia Alcántara será moderadora del XIX congreso…

                            Milton Morrison reafirma alianza con Abinader y anuncia nueva...

                            Milton Morrison reafirma alianza con Abinader y anuncia nueva…

                            Empresarios de Hato Mayor expresan respaldo a Leonel Fernández y fortalecen proyecto político rumbo a 2028

                            Empresarios de Hato Mayor expresan respaldo a Leonel Fernández y fortalecen proyecto político rumbo a 2028

                            PRM en Santo Domingo Norte resalta gestión del presidente...

                            PRM en Santo Domingo Norte resalta gestión del presidente…

                            ARTICULO: De los millones de seguidores al poder: gobernar un país no es hacer un reality en YouTube

                            ARTICULO: De los millones de seguidores al poder: gobernar un país no es hacer un reality en YouTube

                            Estados Unidos no descarta operación militar contra Cuba

                            Estados Unidos no descarta operación militar contra Cuba

                            Tribunal Constitucional ratifica que País Posible es la 7ma fuerza...

                            Tribunal Constitucional ratifica que País Posible es la 7ma fuerza…

                            Sismo en Colombia suma 181 fallecidos

                            Sismo en Colombia suma 181 fallecidos

                            Trending Tags

                            • Deportes
                              • All
                              • Atletas Dominicanos
                              • Béisbol
                              DR Open Kiteboarding Championship reúne atletas de 15 países y reafirma a Cabarete como capital del kitesurf del Caribe

                              Cabarete se corona como capital histórica del kitesurf con el DR Open Championship 2026

                              El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

                              El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

                              La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

                              La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

                              Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

                              Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

                              Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

                              Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

                              Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

                              Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

                              Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

                              Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

                              30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                              Buffalo recibe a Montreal para abrir la segunda ronda

                              Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

                              Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

                              Trending Tags

                              • Economía
                                • All
                                • Combustibles
                                • Energía
                                • Indicadores Económicos
                                • Sector Energético
                                • Turismo
                                Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

                                Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

                                Aventúrate RD 2026

                                Aventúrate RD 2026 revela agenda oficial y consolida el turismo de aventura dominicano

                                WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

                                WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

                                Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

                                Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

                                Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

                                Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

                                Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

                                Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

                                El ministerio de Turismo cierra con éxito festival gastronómico “Saborea el Paraíso” en Sánchez, Samaná

                                El Ministerio de Turismo celebra un exitoso cierre del festival gastronómico «Saborea el Paraíso» en Sánchez, Samaná

                                El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

                                El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

                                Más allá del comercio: los efectos del arancel estadounidense sobre el turismo dominicano

                                Arancel de EE.UU. pone a prueba al turismo dominicano y al silencio oficial del gobierno

                                Trending Tags

                                • Ciencia
                                  • All
                                  • Energía
                                  • Innovación
                                  • Investigación Científica
                                  • Salud y Medicina
                                  • Tecnología Médica
                                  81-year-old admits German cold-case murder of US tourist in 1994

                                  81-year-old admits German cold-case murder of US tourist in 1994

                                  FP podría ganar elecciones en primera vuelta, según Bauta Rojas

                                  FP podría ganar elecciones en primera vuelta, según Bauta Rojas

                                  Lindsay Clancy's trial put postpartrum psychosis on the map. Experts want progress to follow

                                  Lindsay Clancy’s trial put postpartrum psychosis on the map. Experts want progress to follow

                                  Israel opens up tenders for controversial West Bank settlement project

                                  Israel opens up tenders for controversial West Bank settlement project

                                  Target receives $1bn boost from Trump tariff refunds

                                  Target receives $1bn boost from Trump tariff refunds

                                  Ukrainian man arrested in Croatia over Nord Stream blasts

                                  Ukrainian man arrested in Croatia over Nord Stream blasts

                                  Helicopter crashes at Mount Ololokwe in Kenya, killing six people

                                  Helicopter crashes at Mount Ololokwe in Kenya, killing six people

                                  Apparent human remains found in US reservoir as water levels hit record low

                                  Apparent human remains found in US reservoir as water levels hit record low

                                  Erin Patterson appeal: Mushroom murderer's trial undermined by hotel mix-up, court told

                                  Erin Patterson appeal: Mushroom murderer’s trial undermined by hotel mix-up, court told

                                  Trending Tags

                                  • Tecnología
                                    • All
                                    • Aplicaciones
                                    • Inteligencia Artificial
                                    ¿Qué es mejor que ver eclipses desde el espacio? Prueba las luces del norte y del sur.

                                    ¿Qué es mejor que ver eclipses desde el espacio? Prueba las luces del norte y del sur.

                                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                    Un telescopio espía la estrella más rápida de la Vía Láctea girando alrededor de un agujero negro

                                    VentureBeat nombra a Rob Strechay como su primer analista principal, ampliando su impulso de investigación de IA empresarial

                                    VentureBeat nombra a Rob Strechay como primer analista principal, ampliando el esfuerzo de investigación de IA empresarial

                                    VentureBeat nombra a Rob Strechay como su primer analista principal, ampliando su impulso de investigación de IA empresarial

                                    VentureBeat nombra a Rob Strechay como su primer analista principal, ampliando su impulso de investigación de IA empresarial

                                    Desde perros robot hasta ayudantes, China muestra sus ambiciones en materia de robótica en una conferencia

                                    Desde perros robot hasta ayudantes, China muestra sus ambiciones en materia de robótica en una conferencia

                                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                    Fascinantes medusas ocupan un lugar central en el museo de EE. UU. dedicado a estas criaturas

                                    China recupera la etapa de un cohete en tierra por primera vez, después de una recuperación anterior en el mar

                                    China recupera la etapa de un cohete en tierra por primera vez, después de una recuperación anterior en el mar

                                    Un almacén arde cerca de Moscú. En Ucrania, los equipos de drones se preparan para atacar nuevamente

                                    Un almacén arde cerca de Moscú. En Ucrania, los equipos de drones se preparan para atacar nuevamente

                                    Las acciones del fabricante chino de robots humanoides Unitree se disparan en su debut comercial en Shanghai

                                    Las acciones del fabricante chino de robots humanoides Unitree se disparan en su debut comercial en Shanghai

                                    Trending Tags

                                    • Entretenimiento
                                      • All
                                      • Cine y Series
                                      • Cultura Digital
                                      • Cultura Popular
                                      • Gastronomía
                                      • Música
                                      Después de perder a un amigo y escribir 'Say So', Dan + Shay regresan con la autobiográfica 'Young'

                                      Después de perder a un amigo y escribir ‘Say So’, Dan + Shay regresan con la autobiográfica ‘Young’

                                      30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                      El Centro Kennedy dice a la corte que no intentará restaurar el nombre de Trump en el edificio antes del 8 de septiembre

                                      Patólogo forense detalla las heridas fatales de Tupac Shakur en el juicio de Duane 'Keffe D' Davis

                                      Patólogo forense detalla las heridas fatales de Tupac Shakur en el juicio de Duane ‘Keffe D’ Davis

                                      El cofundador de ESPN, Bill Rasmussen, muere a los 93 años por los efectos de la enfermedad de Parkinson

                                      El cofundador de ESPN, Bill Rasmussen, muere a los 93 años por los efectos de la enfermedad de Parkinson

                                      Fox Sports transmitirá 35 partidos de voleibol femenino, incluidos 8 en Fox

                                      Fox Sports transmitirá 35 partidos de voleibol femenino, incluidos 8 en Fox

                                      30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                      Una película de animación china calificada de «terrible» se convierte en un éxito de taquilla

                                      Shakira realiza visita sorpresa a Colombia afectada por el terremoto y se compromete a construir nuevas escuelas

                                      Shakira realiza visita sorpresa a Colombia afectada por el terremoto y se compromete a construir nuevas escuelas

                                      Bonnie Tyler es recordada como estrella mundial en su funeral en Gales

                                      Bonnie Tyler es recordada como estrella mundial en su funeral en Gales

                                      Dave Marsh, biógrafo y crítico musical de Bruce Springsteen, muere a los 76 años

                                      Dave Marsh, biógrafo y crítico musical de Bruce Springsteen, muere a los 76 años

                                      Trending Tags

                                      No Result
                                      View All Result
                                      Despertar Matinal
                                      No Result
                                      View All Result

                                      Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start

                                      by — Redacción Despertar Matinal
                                      23 de julio de 2026
                                      in Tecnología
                                      0
                                      Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start
                                      0
                                      SHARES
                                      11
                                      VIEWS
                                      Share on FacebookShare on Twitter

                                      Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with today’s launch of FLUX 3, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt — and to extend the same underlying architecture to robotic vision and actions.

                                      The Freiburg, Germany-based AI lab says FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface.

                                      That distinction is central to the company’s pitch: BFL wants enterprises to think about creative generation, simulation, computer use and robotics as connected applications of a single capability it calls visual intelligence — models, in the company’s words, «that can perceive, predict, and act across physical and digital environments.» This release marks BFL’s first public video generation model.

                                      FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming, open source FLUX 3 Dev. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a gated «Early Access» program now, to which anyone can apply, but which BFL must approve.

                                      There is presently no public access through BFL’s application programming interface (API) or those of partners yet, but the company says FLUX 3 Image will roll out in the coming weeks, followed by general availability. The limited initial availability rollout echoes the release strategies of new models from other frontier labs in the U.S. lately, including Anthropic and OpenAI, though those were ostensibly for security concerns and due to government request.

                                      What the company has not announced is pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts or any image-model benchmarks at all. Enterprise buyers therefore cannot yet calculate total cost of ownership or independently reproduce the video comparisons.

                                      Another big notable omission: FLUX 3 is not launching with downloadable weights at this time, nor an open source license. BFL says faster and open-weight versions will arrive later this year, and its technical blog names FLUX 3 Dev as «open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction» — a considerably broader commitment than any previous FLUX Dev release, all of which covered images only.

                                      But it arrives last in the sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside — or soon after — a major model announcement will have to wait. That delay does not negate the company’s commitment, but it is disappointing given the role open weights have played in FLUX’s adoption thus far.

                                      Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoption

                                      BFL has published several benchmark comparisons, but they’re qualified as preliminary — with full benchmark results and methodology to be published later during broader general availability.

                                      In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google’s Gemini Omni Flash in 52%.

                                      Flux 3 internal evaluation by viewers of 10-second 720p video generations from various models. Credit: Black Forest Labs

                                      One caveat travels with every one of those figures, and it comes from BFL itself. The chart carrying the results is labeled a «preliminary evaluation of an early FLUX 3 candidate» — meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. That cuts both ways: the shipping model may perform better, but nothing published today measures what customers will actually call.

                                      Luma Ray 3.2 and Runway Gen-4.5, where FLUX 3 posted 93% and 77%, are the softest comparisons on the list — established products, but not the models currently setting the pace in independent video rankings. Those are real wins, and they are the ones least likely to change an enterprise shortlist.

                                      Seedance 2.0, at 52%, is a statistical coin flip against a model most Western enterprises cannot currently procure. ByteDance indefinitely postponed Seedance 2.0’s international rollout after Netflix, Warner Bros., Disney, Paramount and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a frozen product is neither a strong claim nor a damaging one.

                                      Gemini Omni Flash, also at 52%, matters much more. Omni is the closest large-platform analogue to what FLUX 3 is attempting — multimodal input, video and audio-aware creation, conversational editing — and by BFL’s own measurement, the two are indistinguishable on 10-second text-to-video quality.

                                      Google’s advantage in that matchup is that Omni is generally available via Google’s Gemini API for $0.10 per second of generated 720p video, or a 10-second clip for around.

                                      One regional wrinkle matters for a German company’s home market. Editing uploaded video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video the model itself generated is permitted. A European enterprise that wants to run its existing footage through a generative editing pass cannot currently do so on Omni Flash.

                                      Here’s a rough guide for enterprises considering which video models to rely upon:

                                      Model

                                      Max single-generation duration

                                      Max resolution

                                      Key constraints

                                      Price per 10-second clip (720p)

                                      Price per 10-second clip (1080p)

                                      Price per 10-second clip (4K)

                                      FLUX 3 Video

                                      20 seconds

                                      Not stated; evaluations run at 720p

                                      Early access; no published SLA or pricing

                                      Not announced

                                      Not announced

                                      Not announced

                                      HappyHorse 1.1

                                      15 seconds

                                      1080p

                                      No 4K; closed weights

                                      Not published (v1.0 reseller rate is ~$1.82)

                                      Not published (v1.0 reseller rate is ~$3.12)

                                      n/a

                                      Veo 3.1

                                      Per-second billing

                                      4K

                                      Supports clip extension; preview

                                      $4.00

                                      $4.00

                                      $6.00

                                      Veo 3.1 Fast

                                      Per-second billing

                                      4K

                                      Preview

                                      $1.00

                                      $1.20

                                      $3.00

                                      Veo 3.1 Lite

                                      Per-second billing

                                      1080p

                                      No 4K, no clip extension; preview

                                      $0.50

                                      $0.80

                                      n/a

                                      Gemini Omni Flash

                                      10 seconds (3s minimum)

                                      720p at 24 FPS

                                      Preview abd no EU access

                                      $1.00

                                      n/a

                                      n/a

                                      One architecture for media generation and physical action

                                      FLUX 3 builds on Self-Flow, BFL’s method for aligning multimodal understanding and generation within one architecture, publicized back in March 2026.

                                      The company says it significantly scaled up compute and data to train across video, images and audio simultaneously, and that testing showed video generation and action prediction do not require separate foundations — the same architecture could be extended to action prediction without sacrificing what it learned from video.

                                      «We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture,» said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. «True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express.»

                                      He put the case more bluntly elsewhere in the announcement: «You can’t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.»

                                      BFL says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart.

                                      For creative software companies, the appeal is consolidation. A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models.

                                      For robotics teams, the potential value is data efficiency. Models that already encode motion, object behavior and physical change may need less task-specific robot training than systems starting from raw demonstrations.

                                      What FLUX 3 Video can actually do

                                      The video tier is the most concretely specified part of the launch, and it settles a question that had been circulating as rumor: FLUX 3 generates clips of up to 20 seconds with audio in a single generation.

                                      Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio — though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching OpenAI’s discontinued Sora model.

                                      The capability list BFL published covers:

                                      • Text-to-video generation.

                                      • Image-to-video generation, either animating from a starting frame or using images as visual references.

                                      • Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context.

                                      • Generative video-audio continuation from existing video and audio input.

                                      • Keyframe-to-video generation for controlled transitions between defined moments.

                                      • Multilingual dialogue.

                                      • A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics.

                                      • Typography generation and animated design.

                                      • Agentic chaining of individual clips into longer, multi-shot sequences.

                                      That last item is the one enterprise video teams should look at hardest. BFL claims the capabilities combine to produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it addresses the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots.

                                      It is also the capability where competition is most direct. HappyHorse 1.1’s headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to hold identity stable across generated footage — the same problem, approached at the input layer rather than through agentic clip chaining. Alibaba also claims zero-drift lip sync and has specifically targeted the artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening. Character consistency is where this category is being contested, and both companies know it.

                                      BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, the company says preliminary evaluations conducted during midtraining show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. It published no image benchmarks or win rates.

                                      FLUX-mimic tests whether video models can become robot models

                                      BFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access.

                                      The technical blog describes two distinct routes to action prediction: integrating native action prediction directly into FLUX 3, scaling up the initial Self-Flow work; and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic is the second route — the FLUX 3 backbone combined with mimic’s robot-learning and production-deployment expertise in dexterous manipulation.

                                      FLUX-mimic is designed for general-purpose robotic manipulation: helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with far less task-specific data.

                                      BFL and Mimic Robotics say that depending on task difficulty, the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours.

                                      «The hardest part of robotics is data,» said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. «Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning.»

                                      BFL argues that a model trained only on images cannot understand a world that «moves, sounds, changes, and responds,» and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni.

                                      Its developer documentation cites «world knowledge» that combines «an understanding of physics» with Gemini’s grasp of history, science and cultural context. Its marketing is blunter still: «Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different,» the company posted in June, crediting the model with «an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic.»

                                      The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it.

                                      There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of it indirectly. Nothing else on offer captures it at all.

                                      Open weights helped make FLUX an industry standard

                                      BFL officially launched in summer 2024 and gained a name for itself in the AI industry in the intervening two years for its commitment to open sourcing high-quality AI image models beloved by developers, creatives, and enterprises.

                                      The company’s founders, including Rombach, Andreas Blattmann and Patrick Esser, previously helped create VQGAN, latent diffusion and Stable Diffusion, the latter the open source technology that kicked off broad AI generation capabilities for the masses and currently used by many AI image generators and companies.

                                      That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart and Nous Research’s Hermes Agent, among other platforms, and the company cites film director Martin Scorsese among professional users.

                                      Wired magazine described Black Forest Labs as a relatively small company that nevertheless became a leading competitor to Silicon Valley’s largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on AI code sharing community Hugging Face. The company says it now runs a 100-person team across Freiburg and San Francisco.

                                      FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev and related control models, released shortly after the firm’s launch, gave researchers and creative-tool developers access to downloadable checkpoints, local inference and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license.

                                      The company continued that pattern with FLUX.2 Dev in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. Black Forest Labs called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code and optimized implementations for consumer Nvidia GPUs.

                                      FLUX 3 Dev raises the stakes on that evaluation. Previous Dev releases were image models. This one is described as a multimodal backbone spanning video, audio, image and action prediction — meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery. BFL hasn’t yet shared information about its license, the parameter count, quantizations or hardware requirements.

                                      The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and let teams adapt FLUX 3 to their own data, products and workflows.

                                      The financial backing behind FLUX 3 is worth noting alongside the technical claims. Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva and Deutsche Telekom’s T.Capital.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with today’s launch of FLUX 3, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt — and to extend the same underlying architecture to robotic vision and actions.

                                      The Freiburg, Germany-based AI lab says FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface.

                                      That distinction is central to the company’s pitch: BFL wants enterprises to think about creative generation, simulation, computer use and robotics as connected applications of a single capability it calls visual intelligence — models, in the company’s words, «that can perceive, predict, and act across physical and digital environments.» This release marks BFL’s first public video generation model.

                                      FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming, open source FLUX 3 Dev. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a gated «Early Access» program now, to which anyone can apply, but which BFL must approve.

                                      There is presently no public access through BFL’s application programming interface (API) or those of partners yet, but the company says FLUX 3 Image will roll out in the coming weeks, followed by general availability. The limited initial availability rollout echoes the release strategies of new models from other frontier labs in the U.S. lately, including Anthropic and OpenAI, though those were ostensibly for security concerns and due to government request.

                                      What the company has not announced is pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts or any image-model benchmarks at all. Enterprise buyers therefore cannot yet calculate total cost of ownership or independently reproduce the video comparisons.

                                      Another big notable omission: FLUX 3 is not launching with downloadable weights at this time, nor an open source license. BFL says faster and open-weight versions will arrive later this year, and its technical blog names FLUX 3 Dev as «open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction» — a considerably broader commitment than any previous FLUX Dev release, all of which covered images only.

                                      But it arrives last in the sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside — or soon after — a major model announcement will have to wait. That delay does not negate the company’s commitment, but it is disappointing given the role open weights have played in FLUX’s adoption thus far.

                                      Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoption

                                      BFL has published several benchmark comparisons, but they’re qualified as preliminary — with full benchmark results and methodology to be published later during broader general availability.

                                      In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google’s Gemini Omni Flash in 52%.

                                      Flux 3 internal evaluation by viewers of 10-second 720p video generations from various models. Credit: Black Forest Labs

                                      One caveat travels with every one of those figures, and it comes from BFL itself. The chart carrying the results is labeled a «preliminary evaluation of an early FLUX 3 candidate» — meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. That cuts both ways: the shipping model may perform better, but nothing published today measures what customers will actually call.

                                      Luma Ray 3.2 and Runway Gen-4.5, where FLUX 3 posted 93% and 77%, are the softest comparisons on the list — established products, but not the models currently setting the pace in independent video rankings. Those are real wins, and they are the ones least likely to change an enterprise shortlist.

                                      Seedance 2.0, at 52%, is a statistical coin flip against a model most Western enterprises cannot currently procure. ByteDance indefinitely postponed Seedance 2.0’s international rollout after Netflix, Warner Bros., Disney, Paramount and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a frozen product is neither a strong claim nor a damaging one.

                                      Gemini Omni Flash, also at 52%, matters much more. Omni is the closest large-platform analogue to what FLUX 3 is attempting — multimodal input, video and audio-aware creation, conversational editing — and by BFL’s own measurement, the two are indistinguishable on 10-second text-to-video quality.

                                      Google’s advantage in that matchup is that Omni is generally available via Google’s Gemini API for $0.10 per second of generated 720p video, or a 10-second clip for around.

                                      One regional wrinkle matters for a German company’s home market. Editing uploaded video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video the model itself generated is permitted. A European enterprise that wants to run its existing footage through a generative editing pass cannot currently do so on Omni Flash.

                                      Here’s a rough guide for enterprises considering which video models to rely upon:

                                      Model

                                      Max single-generation duration

                                      Max resolution

                                      Key constraints

                                      Price per 10-second clip (720p)

                                      Price per 10-second clip (1080p)

                                      Price per 10-second clip (4K)

                                      FLUX 3 Video

                                      20 seconds

                                      Not stated; evaluations run at 720p

                                      Early access; no published SLA or pricing

                                      Not announced

                                      Not announced

                                      Not announced

                                      HappyHorse 1.1

                                      15 seconds

                                      1080p

                                      No 4K; closed weights

                                      Not published (v1.0 reseller rate is ~$1.82)

                                      Not published (v1.0 reseller rate is ~$3.12)

                                      n/a

                                      Veo 3.1

                                      Per-second billing

                                      4K

                                      Supports clip extension; preview

                                      $4.00

                                      $4.00

                                      $6.00

                                      Veo 3.1 Fast

                                      Per-second billing

                                      4K

                                      Preview

                                      $1.00

                                      $1.20

                                      $3.00

                                      Veo 3.1 Lite

                                      Per-second billing

                                      1080p

                                      No 4K, no clip extension; preview

                                      $0.50

                                      $0.80

                                      n/a

                                      Gemini Omni Flash

                                      10 seconds (3s minimum)

                                      720p at 24 FPS

                                      Preview abd no EU access

                                      $1.00

                                      n/a

                                      n/a

                                      One architecture for media generation and physical action

                                      FLUX 3 builds on Self-Flow, BFL’s method for aligning multimodal understanding and generation within one architecture, publicized back in March 2026.

                                      The company says it significantly scaled up compute and data to train across video, images and audio simultaneously, and that testing showed video generation and action prediction do not require separate foundations — the same architecture could be extended to action prediction without sacrificing what it learned from video.

                                      «We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture,» said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. «True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express.»

                                      He put the case more bluntly elsewhere in the announcement: «You can’t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.»

                                      BFL says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart.

                                      For creative software companies, the appeal is consolidation. A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models.

                                      For robotics teams, the potential value is data efficiency. Models that already encode motion, object behavior and physical change may need less task-specific robot training than systems starting from raw demonstrations.

                                      What FLUX 3 Video can actually do

                                      The video tier is the most concretely specified part of the launch, and it settles a question that had been circulating as rumor: FLUX 3 generates clips of up to 20 seconds with audio in a single generation.

                                      Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio — though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching OpenAI’s discontinued Sora model.

                                      The capability list BFL published covers:

                                      • Text-to-video generation.

                                      • Image-to-video generation, either animating from a starting frame or using images as visual references.

                                      • Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context.

                                      • Generative video-audio continuation from existing video and audio input.

                                      • Keyframe-to-video generation for controlled transitions between defined moments.

                                      • Multilingual dialogue.

                                      • A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics.

                                      • Typography generation and animated design.

                                      • Agentic chaining of individual clips into longer, multi-shot sequences.

                                      That last item is the one enterprise video teams should look at hardest. BFL claims the capabilities combine to produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it addresses the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots.

                                      It is also the capability where competition is most direct. HappyHorse 1.1’s headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to hold identity stable across generated footage — the same problem, approached at the input layer rather than through agentic clip chaining. Alibaba also claims zero-drift lip sync and has specifically targeted the artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening. Character consistency is where this category is being contested, and both companies know it.

                                      BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, the company says preliminary evaluations conducted during midtraining show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. It published no image benchmarks or win rates.

                                      FLUX-mimic tests whether video models can become robot models

                                      BFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access.

                                      The technical blog describes two distinct routes to action prediction: integrating native action prediction directly into FLUX 3, scaling up the initial Self-Flow work; and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic is the second route — the FLUX 3 backbone combined with mimic’s robot-learning and production-deployment expertise in dexterous manipulation.

                                      FLUX-mimic is designed for general-purpose robotic manipulation: helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with far less task-specific data.

                                      BFL and Mimic Robotics say that depending on task difficulty, the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours.

                                      «The hardest part of robotics is data,» said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. «Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning.»

                                      BFL argues that a model trained only on images cannot understand a world that «moves, sounds, changes, and responds,» and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni.

                                      Its developer documentation cites «world knowledge» that combines «an understanding of physics» with Gemini’s grasp of history, science and cultural context. Its marketing is blunter still: «Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different,» the company posted in June, crediting the model with «an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic.»

                                      The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it.

                                      There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of it indirectly. Nothing else on offer captures it at all.

                                      Open weights helped make FLUX an industry standard

                                      BFL officially launched in summer 2024 and gained a name for itself in the AI industry in the intervening two years for its commitment to open sourcing high-quality AI image models beloved by developers, creatives, and enterprises.

                                      The company’s founders, including Rombach, Andreas Blattmann and Patrick Esser, previously helped create VQGAN, latent diffusion and Stable Diffusion, the latter the open source technology that kicked off broad AI generation capabilities for the masses and currently used by many AI image generators and companies.

                                      That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart and Nous Research’s Hermes Agent, among other platforms, and the company cites film director Martin Scorsese among professional users.

                                      Wired magazine described Black Forest Labs as a relatively small company that nevertheless became a leading competitor to Silicon Valley’s largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on AI code sharing community Hugging Face. The company says it now runs a 100-person team across Freiburg and San Francisco.

                                      FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev and related control models, released shortly after the firm’s launch, gave researchers and creative-tool developers access to downloadable checkpoints, local inference and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license.

                                      The company continued that pattern with FLUX.2 Dev in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. Black Forest Labs called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code and optimized implementations for consumer Nvidia GPUs.

                                      FLUX 3 Dev raises the stakes on that evaluation. Previous Dev releases were image models. This one is described as a multimodal backbone spanning video, audio, image and action prediction — meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery. BFL hasn’t yet shared information about its license, the parameter count, quantizations or hardware requirements.

                                      The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and let teams adapt FLUX 3 to their own data, products and workflows.

                                      The financial backing behind FLUX 3 is worth noting alongside the technical claims. Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva and Deutsche Telekom’s T.Capital.

                                      Tours Colombia Todo el año Tours Colombia Todo el año Tours Colombia Todo el año

                                      Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with today’s launch of FLUX 3, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt — and to extend the same underlying architecture to robotic vision and actions.

                                      The Freiburg, Germany-based AI lab says FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface.

                                      That distinction is central to the company’s pitch: BFL wants enterprises to think about creative generation, simulation, computer use and robotics as connected applications of a single capability it calls visual intelligence — models, in the company’s words, «that can perceive, predict, and act across physical and digital environments.» This release marks BFL’s first public video generation model.

                                      FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming, open source FLUX 3 Dev. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a gated «Early Access» program now, to which anyone can apply, but which BFL must approve.

                                      There is presently no public access through BFL’s application programming interface (API) or those of partners yet, but the company says FLUX 3 Image will roll out in the coming weeks, followed by general availability. The limited initial availability rollout echoes the release strategies of new models from other frontier labs in the U.S. lately, including Anthropic and OpenAI, though those were ostensibly for security concerns and due to government request.

                                      What the company has not announced is pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts or any image-model benchmarks at all. Enterprise buyers therefore cannot yet calculate total cost of ownership or independently reproduce the video comparisons.

                                      Another big notable omission: FLUX 3 is not launching with downloadable weights at this time, nor an open source license. BFL says faster and open-weight versions will arrive later this year, and its technical blog names FLUX 3 Dev as «open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction» — a considerably broader commitment than any previous FLUX Dev release, all of which covered images only.

                                      But it arrives last in the sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside — or soon after — a major model announcement will have to wait. That delay does not negate the company’s commitment, but it is disappointing given the role open weights have played in FLUX’s adoption thus far.

                                      Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoption

                                      BFL has published several benchmark comparisons, but they’re qualified as preliminary — with full benchmark results and methodology to be published later during broader general availability.

                                      In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google’s Gemini Omni Flash in 52%.

                                      Flux 3 internal evaluation by viewers of 10-second 720p video generations from various models. Credit: Black Forest Labs

                                      One caveat travels with every one of those figures, and it comes from BFL itself. The chart carrying the results is labeled a «preliminary evaluation of an early FLUX 3 candidate» — meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. That cuts both ways: the shipping model may perform better, but nothing published today measures what customers will actually call.

                                      Luma Ray 3.2 and Runway Gen-4.5, where FLUX 3 posted 93% and 77%, are the softest comparisons on the list — established products, but not the models currently setting the pace in independent video rankings. Those are real wins, and they are the ones least likely to change an enterprise shortlist.

                                      Seedance 2.0, at 52%, is a statistical coin flip against a model most Western enterprises cannot currently procure. ByteDance indefinitely postponed Seedance 2.0’s international rollout after Netflix, Warner Bros., Disney, Paramount and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a frozen product is neither a strong claim nor a damaging one.

                                      Gemini Omni Flash, also at 52%, matters much more. Omni is the closest large-platform analogue to what FLUX 3 is attempting — multimodal input, video and audio-aware creation, conversational editing — and by BFL’s own measurement, the two are indistinguishable on 10-second text-to-video quality.

                                      Google’s advantage in that matchup is that Omni is generally available via Google’s Gemini API for $0.10 per second of generated 720p video, or a 10-second clip for around.

                                      One regional wrinkle matters for a German company’s home market. Editing uploaded video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video the model itself generated is permitted. A European enterprise that wants to run its existing footage through a generative editing pass cannot currently do so on Omni Flash.

                                      Here’s a rough guide for enterprises considering which video models to rely upon:

                                      Model

                                      Max single-generation duration

                                      Max resolution

                                      Key constraints

                                      Price per 10-second clip (720p)

                                      Price per 10-second clip (1080p)

                                      Price per 10-second clip (4K)

                                      FLUX 3 Video

                                      20 seconds

                                      Not stated; evaluations run at 720p

                                      Early access; no published SLA or pricing

                                      Not announced

                                      Not announced

                                      Not announced

                                      HappyHorse 1.1

                                      15 seconds

                                      1080p

                                      No 4K; closed weights

                                      Not published (v1.0 reseller rate is ~$1.82)

                                      Not published (v1.0 reseller rate is ~$3.12)

                                      n/a

                                      Veo 3.1

                                      Per-second billing

                                      4K

                                      Supports clip extension; preview

                                      $4.00

                                      $4.00

                                      $6.00

                                      Veo 3.1 Fast

                                      Per-second billing

                                      4K

                                      Preview

                                      $1.00

                                      $1.20

                                      $3.00

                                      Veo 3.1 Lite

                                      Per-second billing

                                      1080p

                                      No 4K, no clip extension; preview

                                      $0.50

                                      $0.80

                                      n/a

                                      Gemini Omni Flash

                                      10 seconds (3s minimum)

                                      720p at 24 FPS

                                      Preview abd no EU access

                                      $1.00

                                      n/a

                                      n/a

                                      One architecture for media generation and physical action

                                      FLUX 3 builds on Self-Flow, BFL’s method for aligning multimodal understanding and generation within one architecture, publicized back in March 2026.

                                      The company says it significantly scaled up compute and data to train across video, images and audio simultaneously, and that testing showed video generation and action prediction do not require separate foundations — the same architecture could be extended to action prediction without sacrificing what it learned from video.

                                      «We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture,» said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. «True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express.»

                                      He put the case more bluntly elsewhere in the announcement: «You can’t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.»

                                      BFL says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart.

                                      For creative software companies, the appeal is consolidation. A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models.

                                      For robotics teams, the potential value is data efficiency. Models that already encode motion, object behavior and physical change may need less task-specific robot training than systems starting from raw demonstrations.

                                      What FLUX 3 Video can actually do

                                      The video tier is the most concretely specified part of the launch, and it settles a question that had been circulating as rumor: FLUX 3 generates clips of up to 20 seconds with audio in a single generation.

                                      Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio — though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching OpenAI’s discontinued Sora model.

                                      The capability list BFL published covers:

                                      • Text-to-video generation.

                                      • Image-to-video generation, either animating from a starting frame or using images as visual references.

                                      • Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context.

                                      • Generative video-audio continuation from existing video and audio input.

                                      • Keyframe-to-video generation for controlled transitions between defined moments.

                                      • Multilingual dialogue.

                                      • A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics.

                                      • Typography generation and animated design.

                                      • Agentic chaining of individual clips into longer, multi-shot sequences.

                                      That last item is the one enterprise video teams should look at hardest. BFL claims the capabilities combine to produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it addresses the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots.

                                      It is also the capability where competition is most direct. HappyHorse 1.1’s headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to hold identity stable across generated footage — the same problem, approached at the input layer rather than through agentic clip chaining. Alibaba also claims zero-drift lip sync and has specifically targeted the artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening. Character consistency is where this category is being contested, and both companies know it.

                                      BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, the company says preliminary evaluations conducted during midtraining show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. It published no image benchmarks or win rates.

                                      FLUX-mimic tests whether video models can become robot models

                                      BFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access.

                                      The technical blog describes two distinct routes to action prediction: integrating native action prediction directly into FLUX 3, scaling up the initial Self-Flow work; and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic is the second route — the FLUX 3 backbone combined with mimic’s robot-learning and production-deployment expertise in dexterous manipulation.

                                      FLUX-mimic is designed for general-purpose robotic manipulation: helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with far less task-specific data.

                                      BFL and Mimic Robotics say that depending on task difficulty, the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours.

                                      «The hardest part of robotics is data,» said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. «Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning.»

                                      BFL argues that a model trained only on images cannot understand a world that «moves, sounds, changes, and responds,» and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni.

                                      Its developer documentation cites «world knowledge» that combines «an understanding of physics» with Gemini’s grasp of history, science and cultural context. Its marketing is blunter still: «Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different,» the company posted in June, crediting the model with «an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic.»

                                      The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it.

                                      There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of it indirectly. Nothing else on offer captures it at all.

                                      Open weights helped make FLUX an industry standard

                                      BFL officially launched in summer 2024 and gained a name for itself in the AI industry in the intervening two years for its commitment to open sourcing high-quality AI image models beloved by developers, creatives, and enterprises.

                                      The company’s founders, including Rombach, Andreas Blattmann and Patrick Esser, previously helped create VQGAN, latent diffusion and Stable Diffusion, the latter the open source technology that kicked off broad AI generation capabilities for the masses and currently used by many AI image generators and companies.

                                      That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart and Nous Research’s Hermes Agent, among other platforms, and the company cites film director Martin Scorsese among professional users.

                                      Wired magazine described Black Forest Labs as a relatively small company that nevertheless became a leading competitor to Silicon Valley’s largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on AI code sharing community Hugging Face. The company says it now runs a 100-person team across Freiburg and San Francisco.

                                      FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev and related control models, released shortly after the firm’s launch, gave researchers and creative-tool developers access to downloadable checkpoints, local inference and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license.

                                      The company continued that pattern with FLUX.2 Dev in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. Black Forest Labs called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code and optimized implementations for consumer Nvidia GPUs.

                                      FLUX 3 Dev raises the stakes on that evaluation. Previous Dev releases were image models. This one is described as a multimodal backbone spanning video, audio, image and action prediction — meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery. BFL hasn’t yet shared information about its license, the parameter count, quantizations or hardware requirements.

                                      The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and let teams adapt FLUX 3 to their own data, products and workflows.

                                      The financial backing behind FLUX 3 is worth noting alongside the technical claims. Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva and Deutsche Telekom’s T.Capital.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with today’s launch of FLUX 3, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt — and to extend the same underlying architecture to robotic vision and actions.

                                      The Freiburg, Germany-based AI lab says FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface.

                                      That distinction is central to the company’s pitch: BFL wants enterprises to think about creative generation, simulation, computer use and robotics as connected applications of a single capability it calls visual intelligence — models, in the company’s words, «that can perceive, predict, and act across physical and digital environments.» This release marks BFL’s first public video generation model.

                                      FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming, open source FLUX 3 Dev. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a gated «Early Access» program now, to which anyone can apply, but which BFL must approve.

                                      There is presently no public access through BFL’s application programming interface (API) or those of partners yet, but the company says FLUX 3 Image will roll out in the coming weeks, followed by general availability. The limited initial availability rollout echoes the release strategies of new models from other frontier labs in the U.S. lately, including Anthropic and OpenAI, though those were ostensibly for security concerns and due to government request.

                                      What the company has not announced is pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts or any image-model benchmarks at all. Enterprise buyers therefore cannot yet calculate total cost of ownership or independently reproduce the video comparisons.

                                      Another big notable omission: FLUX 3 is not launching with downloadable weights at this time, nor an open source license. BFL says faster and open-weight versions will arrive later this year, and its technical blog names FLUX 3 Dev as «open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction» — a considerably broader commitment than any previous FLUX Dev release, all of which covered images only.

                                      But it arrives last in the sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside — or soon after — a major model announcement will have to wait. That delay does not negate the company’s commitment, but it is disappointing given the role open weights have played in FLUX’s adoption thus far.

                                      Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoption

                                      BFL has published several benchmark comparisons, but they’re qualified as preliminary — with full benchmark results and methodology to be published later during broader general availability.

                                      In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google’s Gemini Omni Flash in 52%.

                                      Flux 3 internal evaluation by viewers of 10-second 720p video generations from various models. Credit: Black Forest Labs

                                      One caveat travels with every one of those figures, and it comes from BFL itself. The chart carrying the results is labeled a «preliminary evaluation of an early FLUX 3 candidate» — meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. That cuts both ways: the shipping model may perform better, but nothing published today measures what customers will actually call.

                                      Luma Ray 3.2 and Runway Gen-4.5, where FLUX 3 posted 93% and 77%, are the softest comparisons on the list — established products, but not the models currently setting the pace in independent video rankings. Those are real wins, and they are the ones least likely to change an enterprise shortlist.

                                      Seedance 2.0, at 52%, is a statistical coin flip against a model most Western enterprises cannot currently procure. ByteDance indefinitely postponed Seedance 2.0’s international rollout after Netflix, Warner Bros., Disney, Paramount and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a frozen product is neither a strong claim nor a damaging one.

                                      Gemini Omni Flash, also at 52%, matters much more. Omni is the closest large-platform analogue to what FLUX 3 is attempting — multimodal input, video and audio-aware creation, conversational editing — and by BFL’s own measurement, the two are indistinguishable on 10-second text-to-video quality.

                                      Google’s advantage in that matchup is that Omni is generally available via Google’s Gemini API for $0.10 per second of generated 720p video, or a 10-second clip for around.

                                      One regional wrinkle matters for a German company’s home market. Editing uploaded video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video the model itself generated is permitted. A European enterprise that wants to run its existing footage through a generative editing pass cannot currently do so on Omni Flash.

                                      Here’s a rough guide for enterprises considering which video models to rely upon:

                                      Model

                                      Max single-generation duration

                                      Max resolution

                                      Key constraints

                                      Price per 10-second clip (720p)

                                      Price per 10-second clip (1080p)

                                      Price per 10-second clip (4K)

                                      FLUX 3 Video

                                      20 seconds

                                      Not stated; evaluations run at 720p

                                      Early access; no published SLA or pricing

                                      Not announced

                                      Not announced

                                      Not announced

                                      HappyHorse 1.1

                                      15 seconds

                                      1080p

                                      No 4K; closed weights

                                      Not published (v1.0 reseller rate is ~$1.82)

                                      Not published (v1.0 reseller rate is ~$3.12)

                                      n/a

                                      Veo 3.1

                                      Per-second billing

                                      4K

                                      Supports clip extension; preview

                                      $4.00

                                      $4.00

                                      $6.00

                                      Veo 3.1 Fast

                                      Per-second billing

                                      4K

                                      Preview

                                      $1.00

                                      $1.20

                                      $3.00

                                      Veo 3.1 Lite

                                      Per-second billing

                                      1080p

                                      No 4K, no clip extension; preview

                                      $0.50

                                      $0.80

                                      n/a

                                      Gemini Omni Flash

                                      10 seconds (3s minimum)

                                      720p at 24 FPS

                                      Preview abd no EU access

                                      $1.00

                                      n/a

                                      n/a

                                      One architecture for media generation and physical action

                                      FLUX 3 builds on Self-Flow, BFL’s method for aligning multimodal understanding and generation within one architecture, publicized back in March 2026.

                                      The company says it significantly scaled up compute and data to train across video, images and audio simultaneously, and that testing showed video generation and action prediction do not require separate foundations — the same architecture could be extended to action prediction without sacrificing what it learned from video.

                                      «We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture,» said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. «True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express.»

                                      He put the case more bluntly elsewhere in the announcement: «You can’t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.»

                                      BFL says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart.

                                      For creative software companies, the appeal is consolidation. A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models.

                                      For robotics teams, the potential value is data efficiency. Models that already encode motion, object behavior and physical change may need less task-specific robot training than systems starting from raw demonstrations.

                                      What FLUX 3 Video can actually do

                                      The video tier is the most concretely specified part of the launch, and it settles a question that had been circulating as rumor: FLUX 3 generates clips of up to 20 seconds with audio in a single generation.

                                      Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio — though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching OpenAI’s discontinued Sora model.

                                      The capability list BFL published covers:

                                      • Text-to-video generation.

                                      • Image-to-video generation, either animating from a starting frame or using images as visual references.

                                      • Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context.

                                      • Generative video-audio continuation from existing video and audio input.

                                      • Keyframe-to-video generation for controlled transitions between defined moments.

                                      • Multilingual dialogue.

                                      • A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics.

                                      • Typography generation and animated design.

                                      • Agentic chaining of individual clips into longer, multi-shot sequences.

                                      That last item is the one enterprise video teams should look at hardest. BFL claims the capabilities combine to produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it addresses the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots.

                                      It is also the capability where competition is most direct. HappyHorse 1.1’s headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to hold identity stable across generated footage — the same problem, approached at the input layer rather than through agentic clip chaining. Alibaba also claims zero-drift lip sync and has specifically targeted the artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening. Character consistency is where this category is being contested, and both companies know it.

                                      BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, the company says preliminary evaluations conducted during midtraining show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. It published no image benchmarks or win rates.

                                      FLUX-mimic tests whether video models can become robot models

                                      BFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access.

                                      The technical blog describes two distinct routes to action prediction: integrating native action prediction directly into FLUX 3, scaling up the initial Self-Flow work; and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic is the second route — the FLUX 3 backbone combined with mimic’s robot-learning and production-deployment expertise in dexterous manipulation.

                                      FLUX-mimic is designed for general-purpose robotic manipulation: helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with far less task-specific data.

                                      BFL and Mimic Robotics say that depending on task difficulty, the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours.

                                      «The hardest part of robotics is data,» said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. «Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning.»

                                      BFL argues that a model trained only on images cannot understand a world that «moves, sounds, changes, and responds,» and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni.

                                      Its developer documentation cites «world knowledge» that combines «an understanding of physics» with Gemini’s grasp of history, science and cultural context. Its marketing is blunter still: «Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different,» the company posted in June, crediting the model with «an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic.»

                                      The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it.

                                      There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of it indirectly. Nothing else on offer captures it at all.

                                      Open weights helped make FLUX an industry standard

                                      BFL officially launched in summer 2024 and gained a name for itself in the AI industry in the intervening two years for its commitment to open sourcing high-quality AI image models beloved by developers, creatives, and enterprises.

                                      The company’s founders, including Rombach, Andreas Blattmann and Patrick Esser, previously helped create VQGAN, latent diffusion and Stable Diffusion, the latter the open source technology that kicked off broad AI generation capabilities for the masses and currently used by many AI image generators and companies.

                                      That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart and Nous Research’s Hermes Agent, among other platforms, and the company cites film director Martin Scorsese among professional users.

                                      Wired magazine described Black Forest Labs as a relatively small company that nevertheless became a leading competitor to Silicon Valley’s largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on AI code sharing community Hugging Face. The company says it now runs a 100-person team across Freiburg and San Francisco.

                                      FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev and related control models, released shortly after the firm’s launch, gave researchers and creative-tool developers access to downloadable checkpoints, local inference and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license.

                                      The company continued that pattern with FLUX.2 Dev in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. Black Forest Labs called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code and optimized implementations for consumer Nvidia GPUs.

                                      FLUX 3 Dev raises the stakes on that evaluation. Previous Dev releases were image models. This one is described as a multimodal backbone spanning video, audio, image and action prediction — meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery. BFL hasn’t yet shared information about its license, the parameter count, quantizations or hardware requirements.

                                      The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and let teams adapt FLUX 3 to their own data, products and workflows.

                                      The financial backing behind FLUX 3 is worth noting alongside the technical claims. Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva and Deutsche Telekom’s T.Capital.

                                      ¡No te pierdas las noticias destacadas!

                                      Suscríbete y recibe las historias más importantes del día.

                                      Al suscribirte aceptas nuestros términos y condiciones y política de privacidad.

                                      Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with today’s launch of FLUX 3, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt — and to extend the same underlying architecture to robotic vision and actions.

                                      The Freiburg, Germany-based AI lab says FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface.

                                      That distinction is central to the company’s pitch: BFL wants enterprises to think about creative generation, simulation, computer use and robotics as connected applications of a single capability it calls visual intelligence — models, in the company’s words, «that can perceive, predict, and act across physical and digital environments.» This release marks BFL’s first public video generation model.

                                      FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming, open source FLUX 3 Dev. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a gated «Early Access» program now, to which anyone can apply, but which BFL must approve.

                                      There is presently no public access through BFL’s application programming interface (API) or those of partners yet, but the company says FLUX 3 Image will roll out in the coming weeks, followed by general availability. The limited initial availability rollout echoes the release strategies of new models from other frontier labs in the U.S. lately, including Anthropic and OpenAI, though those were ostensibly for security concerns and due to government request.

                                      What the company has not announced is pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts or any image-model benchmarks at all. Enterprise buyers therefore cannot yet calculate total cost of ownership or independently reproduce the video comparisons.

                                      Another big notable omission: FLUX 3 is not launching with downloadable weights at this time, nor an open source license. BFL says faster and open-weight versions will arrive later this year, and its technical blog names FLUX 3 Dev as «open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction» — a considerably broader commitment than any previous FLUX Dev release, all of which covered images only.

                                      But it arrives last in the sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside — or soon after — a major model announcement will have to wait. That delay does not negate the company’s commitment, but it is disappointing given the role open weights have played in FLUX’s adoption thus far.

                                      Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoption

                                      BFL has published several benchmark comparisons, but they’re qualified as preliminary — with full benchmark results and methodology to be published later during broader general availability.

                                      In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google’s Gemini Omni Flash in 52%.

                                      Flux 3 internal evaluation by viewers of 10-second 720p video generations from various models. Credit: Black Forest Labs

                                      One caveat travels with every one of those figures, and it comes from BFL itself. The chart carrying the results is labeled a «preliminary evaluation of an early FLUX 3 candidate» — meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. That cuts both ways: the shipping model may perform better, but nothing published today measures what customers will actually call.

                                      Luma Ray 3.2 and Runway Gen-4.5, where FLUX 3 posted 93% and 77%, are the softest comparisons on the list — established products, but not the models currently setting the pace in independent video rankings. Those are real wins, and they are the ones least likely to change an enterprise shortlist.

                                      Seedance 2.0, at 52%, is a statistical coin flip against a model most Western enterprises cannot currently procure. ByteDance indefinitely postponed Seedance 2.0’s international rollout after Netflix, Warner Bros., Disney, Paramount and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a frozen product is neither a strong claim nor a damaging one.

                                      Gemini Omni Flash, also at 52%, matters much more. Omni is the closest large-platform analogue to what FLUX 3 is attempting — multimodal input, video and audio-aware creation, conversational editing — and by BFL’s own measurement, the two are indistinguishable on 10-second text-to-video quality.

                                      Google’s advantage in that matchup is that Omni is generally available via Google’s Gemini API for $0.10 per second of generated 720p video, or a 10-second clip for around.

                                      One regional wrinkle matters for a German company’s home market. Editing uploaded video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video the model itself generated is permitted. A European enterprise that wants to run its existing footage through a generative editing pass cannot currently do so on Omni Flash.

                                      Here’s a rough guide for enterprises considering which video models to rely upon:

                                      Model

                                      Max single-generation duration

                                      Max resolution

                                      Key constraints

                                      Price per 10-second clip (720p)

                                      Price per 10-second clip (1080p)

                                      Price per 10-second clip (4K)

                                      FLUX 3 Video

                                      20 seconds

                                      Not stated; evaluations run at 720p

                                      Early access; no published SLA or pricing

                                      Not announced

                                      Not announced

                                      Not announced

                                      HappyHorse 1.1

                                      15 seconds

                                      1080p

                                      No 4K; closed weights

                                      Not published (v1.0 reseller rate is ~$1.82)

                                      Not published (v1.0 reseller rate is ~$3.12)

                                      n/a

                                      Veo 3.1

                                      Per-second billing

                                      4K

                                      Supports clip extension; preview

                                      $4.00

                                      $4.00

                                      $6.00

                                      Veo 3.1 Fast

                                      Per-second billing

                                      4K

                                      Preview

                                      $1.00

                                      $1.20

                                      $3.00

                                      Veo 3.1 Lite

                                      Per-second billing

                                      1080p

                                      No 4K, no clip extension; preview

                                      $0.50

                                      $0.80

                                      n/a

                                      Gemini Omni Flash

                                      10 seconds (3s minimum)

                                      720p at 24 FPS

                                      Preview abd no EU access

                                      $1.00

                                      n/a

                                      n/a

                                      One architecture for media generation and physical action

                                      FLUX 3 builds on Self-Flow, BFL’s method for aligning multimodal understanding and generation within one architecture, publicized back in March 2026.

                                      The company says it significantly scaled up compute and data to train across video, images and audio simultaneously, and that testing showed video generation and action prediction do not require separate foundations — the same architecture could be extended to action prediction without sacrificing what it learned from video.

                                      «We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture,» said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. «True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express.»

                                      He put the case more bluntly elsewhere in the announcement: «You can’t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.»

                                      BFL says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart.

                                      For creative software companies, the appeal is consolidation. A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models.

                                      For robotics teams, the potential value is data efficiency. Models that already encode motion, object behavior and physical change may need less task-specific robot training than systems starting from raw demonstrations.

                                      What FLUX 3 Video can actually do

                                      The video tier is the most concretely specified part of the launch, and it settles a question that had been circulating as rumor: FLUX 3 generates clips of up to 20 seconds with audio in a single generation.

                                      Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio — though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching OpenAI’s discontinued Sora model.

                                      The capability list BFL published covers:

                                      • Text-to-video generation.

                                      • Image-to-video generation, either animating from a starting frame or using images as visual references.

                                      • Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context.

                                      • Generative video-audio continuation from existing video and audio input.

                                      • Keyframe-to-video generation for controlled transitions between defined moments.

                                      • Multilingual dialogue.

                                      • A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics.

                                      • Typography generation and animated design.

                                      • Agentic chaining of individual clips into longer, multi-shot sequences.

                                      That last item is the one enterprise video teams should look at hardest. BFL claims the capabilities combine to produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it addresses the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots.

                                      It is also the capability where competition is most direct. HappyHorse 1.1’s headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to hold identity stable across generated footage — the same problem, approached at the input layer rather than through agentic clip chaining. Alibaba also claims zero-drift lip sync and has specifically targeted the artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening. Character consistency is where this category is being contested, and both companies know it.

                                      BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, the company says preliminary evaluations conducted during midtraining show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. It published no image benchmarks or win rates.

                                      FLUX-mimic tests whether video models can become robot models

                                      BFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access.

                                      The technical blog describes two distinct routes to action prediction: integrating native action prediction directly into FLUX 3, scaling up the initial Self-Flow work; and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic is the second route — the FLUX 3 backbone combined with mimic’s robot-learning and production-deployment expertise in dexterous manipulation.

                                      FLUX-mimic is designed for general-purpose robotic manipulation: helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with far less task-specific data.

                                      BFL and Mimic Robotics say that depending on task difficulty, the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours.

                                      «The hardest part of robotics is data,» said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. «Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning.»

                                      BFL argues that a model trained only on images cannot understand a world that «moves, sounds, changes, and responds,» and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni.

                                      Its developer documentation cites «world knowledge» that combines «an understanding of physics» with Gemini’s grasp of history, science and cultural context. Its marketing is blunter still: «Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different,» the company posted in June, crediting the model with «an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic.»

                                      The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it.

                                      There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of it indirectly. Nothing else on offer captures it at all.

                                      Open weights helped make FLUX an industry standard

                                      BFL officially launched in summer 2024 and gained a name for itself in the AI industry in the intervening two years for its commitment to open sourcing high-quality AI image models beloved by developers, creatives, and enterprises.

                                      The company’s founders, including Rombach, Andreas Blattmann and Patrick Esser, previously helped create VQGAN, latent diffusion and Stable Diffusion, the latter the open source technology that kicked off broad AI generation capabilities for the masses and currently used by many AI image generators and companies.

                                      That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart and Nous Research’s Hermes Agent, among other platforms, and the company cites film director Martin Scorsese among professional users.

                                      Wired magazine described Black Forest Labs as a relatively small company that nevertheless became a leading competitor to Silicon Valley’s largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on AI code sharing community Hugging Face. The company says it now runs a 100-person team across Freiburg and San Francisco.

                                      FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev and related control models, released shortly after the firm’s launch, gave researchers and creative-tool developers access to downloadable checkpoints, local inference and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license.

                                      The company continued that pattern with FLUX.2 Dev in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. Black Forest Labs called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code and optimized implementations for consumer Nvidia GPUs.

                                      FLUX 3 Dev raises the stakes on that evaluation. Previous Dev releases were image models. This one is described as a multimodal backbone spanning video, audio, image and action prediction — meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery. BFL hasn’t yet shared information about its license, the parameter count, quantizations or hardware requirements.

                                      The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and let teams adapt FLUX 3 to their own data, products and workflows.

                                      The financial backing behind FLUX 3 is worth noting alongside the technical claims. Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva and Deutsche Telekom’s T.Capital.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with today’s launch of FLUX 3, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt — and to extend the same underlying architecture to robotic vision and actions.

                                      The Freiburg, Germany-based AI lab says FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface.

                                      That distinction is central to the company’s pitch: BFL wants enterprises to think about creative generation, simulation, computer use and robotics as connected applications of a single capability it calls visual intelligence — models, in the company’s words, «that can perceive, predict, and act across physical and digital environments.» This release marks BFL’s first public video generation model.

                                      FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming, open source FLUX 3 Dev. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a gated «Early Access» program now, to which anyone can apply, but which BFL must approve.

                                      There is presently no public access through BFL’s application programming interface (API) or those of partners yet, but the company says FLUX 3 Image will roll out in the coming weeks, followed by general availability. The limited initial availability rollout echoes the release strategies of new models from other frontier labs in the U.S. lately, including Anthropic and OpenAI, though those were ostensibly for security concerns and due to government request.

                                      What the company has not announced is pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts or any image-model benchmarks at all. Enterprise buyers therefore cannot yet calculate total cost of ownership or independently reproduce the video comparisons.

                                      Another big notable omission: FLUX 3 is not launching with downloadable weights at this time, nor an open source license. BFL says faster and open-weight versions will arrive later this year, and its technical blog names FLUX 3 Dev as «open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction» — a considerably broader commitment than any previous FLUX Dev release, all of which covered images only.

                                      But it arrives last in the sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside — or soon after — a major model announcement will have to wait. That delay does not negate the company’s commitment, but it is disappointing given the role open weights have played in FLUX’s adoption thus far.

                                      Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoption

                                      BFL has published several benchmark comparisons, but they’re qualified as preliminary — with full benchmark results and methodology to be published later during broader general availability.

                                      In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google’s Gemini Omni Flash in 52%.

                                      Flux 3 internal evaluation by viewers of 10-second 720p video generations from various models. Credit: Black Forest Labs

                                      One caveat travels with every one of those figures, and it comes from BFL itself. The chart carrying the results is labeled a «preliminary evaluation of an early FLUX 3 candidate» — meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. That cuts both ways: the shipping model may perform better, but nothing published today measures what customers will actually call.

                                      Luma Ray 3.2 and Runway Gen-4.5, where FLUX 3 posted 93% and 77%, are the softest comparisons on the list — established products, but not the models currently setting the pace in independent video rankings. Those are real wins, and they are the ones least likely to change an enterprise shortlist.

                                      Seedance 2.0, at 52%, is a statistical coin flip against a model most Western enterprises cannot currently procure. ByteDance indefinitely postponed Seedance 2.0’s international rollout after Netflix, Warner Bros., Disney, Paramount and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a frozen product is neither a strong claim nor a damaging one.

                                      Gemini Omni Flash, also at 52%, matters much more. Omni is the closest large-platform analogue to what FLUX 3 is attempting — multimodal input, video and audio-aware creation, conversational editing — and by BFL’s own measurement, the two are indistinguishable on 10-second text-to-video quality.

                                      Google’s advantage in that matchup is that Omni is generally available via Google’s Gemini API for $0.10 per second of generated 720p video, or a 10-second clip for around.

                                      One regional wrinkle matters for a German company’s home market. Editing uploaded video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video the model itself generated is permitted. A European enterprise that wants to run its existing footage through a generative editing pass cannot currently do so on Omni Flash.

                                      Here’s a rough guide for enterprises considering which video models to rely upon:

                                      Model

                                      Max single-generation duration

                                      Max resolution

                                      Key constraints

                                      Price per 10-second clip (720p)

                                      Price per 10-second clip (1080p)

                                      Price per 10-second clip (4K)

                                      FLUX 3 Video

                                      20 seconds

                                      Not stated; evaluations run at 720p

                                      Early access; no published SLA or pricing

                                      Not announced

                                      Not announced

                                      Not announced

                                      HappyHorse 1.1

                                      15 seconds

                                      1080p

                                      No 4K; closed weights

                                      Not published (v1.0 reseller rate is ~$1.82)

                                      Not published (v1.0 reseller rate is ~$3.12)

                                      n/a

                                      Veo 3.1

                                      Per-second billing

                                      4K

                                      Supports clip extension; preview

                                      $4.00

                                      $4.00

                                      $6.00

                                      Veo 3.1 Fast

                                      Per-second billing

                                      4K

                                      Preview

                                      $1.00

                                      $1.20

                                      $3.00

                                      Veo 3.1 Lite

                                      Per-second billing

                                      1080p

                                      No 4K, no clip extension; preview

                                      $0.50

                                      $0.80

                                      n/a

                                      Gemini Omni Flash

                                      10 seconds (3s minimum)

                                      720p at 24 FPS

                                      Preview abd no EU access

                                      $1.00

                                      n/a

                                      n/a

                                      One architecture for media generation and physical action

                                      FLUX 3 builds on Self-Flow, BFL’s method for aligning multimodal understanding and generation within one architecture, publicized back in March 2026.

                                      The company says it significantly scaled up compute and data to train across video, images and audio simultaneously, and that testing showed video generation and action prediction do not require separate foundations — the same architecture could be extended to action prediction without sacrificing what it learned from video.

                                      «We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture,» said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. «True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express.»

                                      He put the case more bluntly elsewhere in the announcement: «You can’t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.»

                                      BFL says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart.

                                      For creative software companies, the appeal is consolidation. A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models.

                                      For robotics teams, the potential value is data efficiency. Models that already encode motion, object behavior and physical change may need less task-specific robot training than systems starting from raw demonstrations.

                                      What FLUX 3 Video can actually do

                                      The video tier is the most concretely specified part of the launch, and it settles a question that had been circulating as rumor: FLUX 3 generates clips of up to 20 seconds with audio in a single generation.

                                      Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio — though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching OpenAI’s discontinued Sora model.

                                      The capability list BFL published covers:

                                      • Text-to-video generation.

                                      • Image-to-video generation, either animating from a starting frame or using images as visual references.

                                      • Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context.

                                      • Generative video-audio continuation from existing video and audio input.

                                      • Keyframe-to-video generation for controlled transitions between defined moments.

                                      • Multilingual dialogue.

                                      • A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics.

                                      • Typography generation and animated design.

                                      • Agentic chaining of individual clips into longer, multi-shot sequences.

                                      That last item is the one enterprise video teams should look at hardest. BFL claims the capabilities combine to produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it addresses the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots.

                                      It is also the capability where competition is most direct. HappyHorse 1.1’s headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to hold identity stable across generated footage — the same problem, approached at the input layer rather than through agentic clip chaining. Alibaba also claims zero-drift lip sync and has specifically targeted the artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening. Character consistency is where this category is being contested, and both companies know it.

                                      BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, the company says preliminary evaluations conducted during midtraining show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. It published no image benchmarks or win rates.

                                      FLUX-mimic tests whether video models can become robot models

                                      BFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access.

                                      The technical blog describes two distinct routes to action prediction: integrating native action prediction directly into FLUX 3, scaling up the initial Self-Flow work; and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic is the second route — the FLUX 3 backbone combined with mimic’s robot-learning and production-deployment expertise in dexterous manipulation.

                                      FLUX-mimic is designed for general-purpose robotic manipulation: helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with far less task-specific data.

                                      BFL and Mimic Robotics say that depending on task difficulty, the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours.

                                      «The hardest part of robotics is data,» said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. «Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning.»

                                      BFL argues that a model trained only on images cannot understand a world that «moves, sounds, changes, and responds,» and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni.

                                      Its developer documentation cites «world knowledge» that combines «an understanding of physics» with Gemini’s grasp of history, science and cultural context. Its marketing is blunter still: «Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different,» the company posted in June, crediting the model with «an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic.»

                                      The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it.

                                      There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of it indirectly. Nothing else on offer captures it at all.

                                      Open weights helped make FLUX an industry standard

                                      BFL officially launched in summer 2024 and gained a name for itself in the AI industry in the intervening two years for its commitment to open sourcing high-quality AI image models beloved by developers, creatives, and enterprises.

                                      The company’s founders, including Rombach, Andreas Blattmann and Patrick Esser, previously helped create VQGAN, latent diffusion and Stable Diffusion, the latter the open source technology that kicked off broad AI generation capabilities for the masses and currently used by many AI image generators and companies.

                                      That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart and Nous Research’s Hermes Agent, among other platforms, and the company cites film director Martin Scorsese among professional users.

                                      Wired magazine described Black Forest Labs as a relatively small company that nevertheless became a leading competitor to Silicon Valley’s largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on AI code sharing community Hugging Face. The company says it now runs a 100-person team across Freiburg and San Francisco.

                                      FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev and related control models, released shortly after the firm’s launch, gave researchers and creative-tool developers access to downloadable checkpoints, local inference and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license.

                                      The company continued that pattern with FLUX.2 Dev in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. Black Forest Labs called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code and optimized implementations for consumer Nvidia GPUs.

                                      FLUX 3 Dev raises the stakes on that evaluation. Previous Dev releases were image models. This one is described as a multimodal backbone spanning video, audio, image and action prediction — meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery. BFL hasn’t yet shared information about its license, the parameter count, quantizations or hardware requirements.

                                      The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and let teams adapt FLUX 3 to their own data, products and workflows.

                                      The financial backing behind FLUX 3 is worth noting alongside the technical claims. Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva and Deutsche Telekom’s T.Capital.

                                      Tours Colombia Todo el año Tours Colombia Todo el año Tours Colombia Todo el año

                                      Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with today’s launch of FLUX 3, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt — and to extend the same underlying architecture to robotic vision and actions.

                                      The Freiburg, Germany-based AI lab says FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface.

                                      That distinction is central to the company’s pitch: BFL wants enterprises to think about creative generation, simulation, computer use and robotics as connected applications of a single capability it calls visual intelligence — models, in the company’s words, «that can perceive, predict, and act across physical and digital environments.» This release marks BFL’s first public video generation model.

                                      FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming, open source FLUX 3 Dev. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a gated «Early Access» program now, to which anyone can apply, but which BFL must approve.

                                      There is presently no public access through BFL’s application programming interface (API) or those of partners yet, but the company says FLUX 3 Image will roll out in the coming weeks, followed by general availability. The limited initial availability rollout echoes the release strategies of new models from other frontier labs in the U.S. lately, including Anthropic and OpenAI, though those were ostensibly for security concerns and due to government request.

                                      What the company has not announced is pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts or any image-model benchmarks at all. Enterprise buyers therefore cannot yet calculate total cost of ownership or independently reproduce the video comparisons.

                                      Another big notable omission: FLUX 3 is not launching with downloadable weights at this time, nor an open source license. BFL says faster and open-weight versions will arrive later this year, and its technical blog names FLUX 3 Dev as «open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction» — a considerably broader commitment than any previous FLUX Dev release, all of which covered images only.

                                      But it arrives last in the sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside — or soon after — a major model announcement will have to wait. That delay does not negate the company’s commitment, but it is disappointing given the role open weights have played in FLUX’s adoption thus far.

                                      Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoption

                                      BFL has published several benchmark comparisons, but they’re qualified as preliminary — with full benchmark results and methodology to be published later during broader general availability.

                                      In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google’s Gemini Omni Flash in 52%.

                                      Flux 3 internal evaluation by viewers of 10-second 720p video generations from various models. Credit: Black Forest Labs

                                      One caveat travels with every one of those figures, and it comes from BFL itself. The chart carrying the results is labeled a «preliminary evaluation of an early FLUX 3 candidate» — meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. That cuts both ways: the shipping model may perform better, but nothing published today measures what customers will actually call.

                                      Luma Ray 3.2 and Runway Gen-4.5, where FLUX 3 posted 93% and 77%, are the softest comparisons on the list — established products, but not the models currently setting the pace in independent video rankings. Those are real wins, and they are the ones least likely to change an enterprise shortlist.

                                      Seedance 2.0, at 52%, is a statistical coin flip against a model most Western enterprises cannot currently procure. ByteDance indefinitely postponed Seedance 2.0’s international rollout after Netflix, Warner Bros., Disney, Paramount and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a frozen product is neither a strong claim nor a damaging one.

                                      Gemini Omni Flash, also at 52%, matters much more. Omni is the closest large-platform analogue to what FLUX 3 is attempting — multimodal input, video and audio-aware creation, conversational editing — and by BFL’s own measurement, the two are indistinguishable on 10-second text-to-video quality.

                                      Google’s advantage in that matchup is that Omni is generally available via Google’s Gemini API for $0.10 per second of generated 720p video, or a 10-second clip for around.

                                      One regional wrinkle matters for a German company’s home market. Editing uploaded video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video the model itself generated is permitted. A European enterprise that wants to run its existing footage through a generative editing pass cannot currently do so on Omni Flash.

                                      Here’s a rough guide for enterprises considering which video models to rely upon:

                                      Model

                                      Max single-generation duration

                                      Max resolution

                                      Key constraints

                                      Price per 10-second clip (720p)

                                      Price per 10-second clip (1080p)

                                      Price per 10-second clip (4K)

                                      FLUX 3 Video

                                      20 seconds

                                      Not stated; evaluations run at 720p

                                      Early access; no published SLA or pricing

                                      Not announced

                                      Not announced

                                      Not announced

                                      HappyHorse 1.1

                                      15 seconds

                                      1080p

                                      No 4K; closed weights

                                      Not published (v1.0 reseller rate is ~$1.82)

                                      Not published (v1.0 reseller rate is ~$3.12)

                                      n/a

                                      Veo 3.1

                                      Per-second billing

                                      4K

                                      Supports clip extension; preview

                                      $4.00

                                      $4.00

                                      $6.00

                                      Veo 3.1 Fast

                                      Per-second billing

                                      4K

                                      Preview

                                      $1.00

                                      $1.20

                                      $3.00

                                      Veo 3.1 Lite

                                      Per-second billing

                                      1080p

                                      No 4K, no clip extension; preview

                                      $0.50

                                      $0.80

                                      n/a

                                      Gemini Omni Flash

                                      10 seconds (3s minimum)

                                      720p at 24 FPS

                                      Preview abd no EU access

                                      $1.00

                                      n/a

                                      n/a

                                      One architecture for media generation and physical action

                                      FLUX 3 builds on Self-Flow, BFL’s method for aligning multimodal understanding and generation within one architecture, publicized back in March 2026.

                                      The company says it significantly scaled up compute and data to train across video, images and audio simultaneously, and that testing showed video generation and action prediction do not require separate foundations — the same architecture could be extended to action prediction without sacrificing what it learned from video.

                                      «We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture,» said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. «True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express.»

                                      He put the case more bluntly elsewhere in the announcement: «You can’t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.»

                                      BFL says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart.

                                      For creative software companies, the appeal is consolidation. A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models.

                                      For robotics teams, the potential value is data efficiency. Models that already encode motion, object behavior and physical change may need less task-specific robot training than systems starting from raw demonstrations.

                                      What FLUX 3 Video can actually do

                                      The video tier is the most concretely specified part of the launch, and it settles a question that had been circulating as rumor: FLUX 3 generates clips of up to 20 seconds with audio in a single generation.

                                      Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio — though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching OpenAI’s discontinued Sora model.

                                      The capability list BFL published covers:

                                      • Text-to-video generation.

                                      • Image-to-video generation, either animating from a starting frame or using images as visual references.

                                      • Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context.

                                      • Generative video-audio continuation from existing video and audio input.

                                      • Keyframe-to-video generation for controlled transitions between defined moments.

                                      • Multilingual dialogue.

                                      • A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics.

                                      • Typography generation and animated design.

                                      • Agentic chaining of individual clips into longer, multi-shot sequences.

                                      That last item is the one enterprise video teams should look at hardest. BFL claims the capabilities combine to produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it addresses the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots.

                                      It is also the capability where competition is most direct. HappyHorse 1.1’s headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to hold identity stable across generated footage — the same problem, approached at the input layer rather than through agentic clip chaining. Alibaba also claims zero-drift lip sync and has specifically targeted the artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening. Character consistency is where this category is being contested, and both companies know it.

                                      BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, the company says preliminary evaluations conducted during midtraining show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. It published no image benchmarks or win rates.

                                      FLUX-mimic tests whether video models can become robot models

                                      BFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access.

                                      The technical blog describes two distinct routes to action prediction: integrating native action prediction directly into FLUX 3, scaling up the initial Self-Flow work; and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic is the second route — the FLUX 3 backbone combined with mimic’s robot-learning and production-deployment expertise in dexterous manipulation.

                                      FLUX-mimic is designed for general-purpose robotic manipulation: helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with far less task-specific data.

                                      BFL and Mimic Robotics say that depending on task difficulty, the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours.

                                      «The hardest part of robotics is data,» said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. «Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning.»

                                      BFL argues that a model trained only on images cannot understand a world that «moves, sounds, changes, and responds,» and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni.

                                      Its developer documentation cites «world knowledge» that combines «an understanding of physics» with Gemini’s grasp of history, science and cultural context. Its marketing is blunter still: «Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different,» the company posted in June, crediting the model with «an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic.»

                                      The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it.

                                      There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of it indirectly. Nothing else on offer captures it at all.

                                      Open weights helped make FLUX an industry standard

                                      BFL officially launched in summer 2024 and gained a name for itself in the AI industry in the intervening two years for its commitment to open sourcing high-quality AI image models beloved by developers, creatives, and enterprises.

                                      The company’s founders, including Rombach, Andreas Blattmann and Patrick Esser, previously helped create VQGAN, latent diffusion and Stable Diffusion, the latter the open source technology that kicked off broad AI generation capabilities for the masses and currently used by many AI image generators and companies.

                                      That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart and Nous Research’s Hermes Agent, among other platforms, and the company cites film director Martin Scorsese among professional users.

                                      Wired magazine described Black Forest Labs as a relatively small company that nevertheless became a leading competitor to Silicon Valley’s largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on AI code sharing community Hugging Face. The company says it now runs a 100-person team across Freiburg and San Francisco.

                                      FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev and related control models, released shortly after the firm’s launch, gave researchers and creative-tool developers access to downloadable checkpoints, local inference and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license.

                                      The company continued that pattern with FLUX.2 Dev in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. Black Forest Labs called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code and optimized implementations for consumer Nvidia GPUs.

                                      FLUX 3 Dev raises the stakes on that evaluation. Previous Dev releases were image models. This one is described as a multimodal backbone spanning video, audio, image and action prediction — meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery. BFL hasn’t yet shared information about its license, the parameter count, quantizations or hardware requirements.

                                      The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and let teams adapt FLUX 3 to their own data, products and workflows.

                                      The financial backing behind FLUX 3 is worth noting alongside the technical claims. Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva and Deutsche Telekom’s T.Capital.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with today’s launch of FLUX 3, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt — and to extend the same underlying architecture to robotic vision and actions.

                                      The Freiburg, Germany-based AI lab says FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface.

                                      That distinction is central to the company’s pitch: BFL wants enterprises to think about creative generation, simulation, computer use and robotics as connected applications of a single capability it calls visual intelligence — models, in the company’s words, «that can perceive, predict, and act across physical and digital environments.» This release marks BFL’s first public video generation model.

                                      FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming, open source FLUX 3 Dev. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a gated «Early Access» program now, to which anyone can apply, but which BFL must approve.

                                      There is presently no public access through BFL’s application programming interface (API) or those of partners yet, but the company says FLUX 3 Image will roll out in the coming weeks, followed by general availability. The limited initial availability rollout echoes the release strategies of new models from other frontier labs in the U.S. lately, including Anthropic and OpenAI, though those were ostensibly for security concerns and due to government request.

                                      What the company has not announced is pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts or any image-model benchmarks at all. Enterprise buyers therefore cannot yet calculate total cost of ownership or independently reproduce the video comparisons.

                                      Another big notable omission: FLUX 3 is not launching with downloadable weights at this time, nor an open source license. BFL says faster and open-weight versions will arrive later this year, and its technical blog names FLUX 3 Dev as «open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction» — a considerably broader commitment than any previous FLUX Dev release, all of which covered images only.

                                      But it arrives last in the sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside — or soon after — a major model announcement will have to wait. That delay does not negate the company’s commitment, but it is disappointing given the role open weights have played in FLUX’s adoption thus far.

                                      Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoption

                                      BFL has published several benchmark comparisons, but they’re qualified as preliminary — with full benchmark results and methodology to be published later during broader general availability.

                                      In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google’s Gemini Omni Flash in 52%.

                                      Flux 3 internal evaluation by viewers of 10-second 720p video generations from various models. Credit: Black Forest Labs

                                      One caveat travels with every one of those figures, and it comes from BFL itself. The chart carrying the results is labeled a «preliminary evaluation of an early FLUX 3 candidate» — meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. That cuts both ways: the shipping model may perform better, but nothing published today measures what customers will actually call.

                                      Luma Ray 3.2 and Runway Gen-4.5, where FLUX 3 posted 93% and 77%, are the softest comparisons on the list — established products, but not the models currently setting the pace in independent video rankings. Those are real wins, and they are the ones least likely to change an enterprise shortlist.

                                      Seedance 2.0, at 52%, is a statistical coin flip against a model most Western enterprises cannot currently procure. ByteDance indefinitely postponed Seedance 2.0’s international rollout after Netflix, Warner Bros., Disney, Paramount and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a frozen product is neither a strong claim nor a damaging one.

                                      Gemini Omni Flash, also at 52%, matters much more. Omni is the closest large-platform analogue to what FLUX 3 is attempting — multimodal input, video and audio-aware creation, conversational editing — and by BFL’s own measurement, the two are indistinguishable on 10-second text-to-video quality.

                                      Google’s advantage in that matchup is that Omni is generally available via Google’s Gemini API for $0.10 per second of generated 720p video, or a 10-second clip for around.

                                      One regional wrinkle matters for a German company’s home market. Editing uploaded video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video the model itself generated is permitted. A European enterprise that wants to run its existing footage through a generative editing pass cannot currently do so on Omni Flash.

                                      Here’s a rough guide for enterprises considering which video models to rely upon:

                                      Model

                                      Max single-generation duration

                                      Max resolution

                                      Key constraints

                                      Price per 10-second clip (720p)

                                      Price per 10-second clip (1080p)

                                      Price per 10-second clip (4K)

                                      FLUX 3 Video

                                      20 seconds

                                      Not stated; evaluations run at 720p

                                      Early access; no published SLA or pricing

                                      Not announced

                                      Not announced

                                      Not announced

                                      HappyHorse 1.1

                                      15 seconds

                                      1080p

                                      No 4K; closed weights

                                      Not published (v1.0 reseller rate is ~$1.82)

                                      Not published (v1.0 reseller rate is ~$3.12)

                                      n/a

                                      Veo 3.1

                                      Per-second billing

                                      4K

                                      Supports clip extension; preview

                                      $4.00

                                      $4.00

                                      $6.00

                                      Veo 3.1 Fast

                                      Per-second billing

                                      4K

                                      Preview

                                      $1.00

                                      $1.20

                                      $3.00

                                      Veo 3.1 Lite

                                      Per-second billing

                                      1080p

                                      No 4K, no clip extension; preview

                                      $0.50

                                      $0.80

                                      n/a

                                      Gemini Omni Flash

                                      10 seconds (3s minimum)

                                      720p at 24 FPS

                                      Preview abd no EU access

                                      $1.00

                                      n/a

                                      n/a

                                      One architecture for media generation and physical action

                                      FLUX 3 builds on Self-Flow, BFL’s method for aligning multimodal understanding and generation within one architecture, publicized back in March 2026.

                                      The company says it significantly scaled up compute and data to train across video, images and audio simultaneously, and that testing showed video generation and action prediction do not require separate foundations — the same architecture could be extended to action prediction without sacrificing what it learned from video.

                                      «We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture,» said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. «True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express.»

                                      He put the case more bluntly elsewhere in the announcement: «You can’t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.»

                                      BFL says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart.

                                      For creative software companies, the appeal is consolidation. A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models.

                                      For robotics teams, the potential value is data efficiency. Models that already encode motion, object behavior and physical change may need less task-specific robot training than systems starting from raw demonstrations.

                                      What FLUX 3 Video can actually do

                                      The video tier is the most concretely specified part of the launch, and it settles a question that had been circulating as rumor: FLUX 3 generates clips of up to 20 seconds with audio in a single generation.

                                      Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio — though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching OpenAI’s discontinued Sora model.

                                      The capability list BFL published covers:

                                      • Text-to-video generation.

                                      • Image-to-video generation, either animating from a starting frame or using images as visual references.

                                      • Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context.

                                      • Generative video-audio continuation from existing video and audio input.

                                      • Keyframe-to-video generation for controlled transitions between defined moments.

                                      • Multilingual dialogue.

                                      • A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics.

                                      • Typography generation and animated design.

                                      • Agentic chaining of individual clips into longer, multi-shot sequences.

                                      That last item is the one enterprise video teams should look at hardest. BFL claims the capabilities combine to produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it addresses the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots.

                                      It is also the capability where competition is most direct. HappyHorse 1.1’s headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to hold identity stable across generated footage — the same problem, approached at the input layer rather than through agentic clip chaining. Alibaba also claims zero-drift lip sync and has specifically targeted the artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening. Character consistency is where this category is being contested, and both companies know it.

                                      BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, the company says preliminary evaluations conducted during midtraining show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. It published no image benchmarks or win rates.

                                      FLUX-mimic tests whether video models can become robot models

                                      BFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access.

                                      The technical blog describes two distinct routes to action prediction: integrating native action prediction directly into FLUX 3, scaling up the initial Self-Flow work; and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic is the second route — the FLUX 3 backbone combined with mimic’s robot-learning and production-deployment expertise in dexterous manipulation.

                                      FLUX-mimic is designed for general-purpose robotic manipulation: helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with far less task-specific data.

                                      BFL and Mimic Robotics say that depending on task difficulty, the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours.

                                      «The hardest part of robotics is data,» said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. «Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning.»

                                      BFL argues that a model trained only on images cannot understand a world that «moves, sounds, changes, and responds,» and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni.

                                      Its developer documentation cites «world knowledge» that combines «an understanding of physics» with Gemini’s grasp of history, science and cultural context. Its marketing is blunter still: «Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different,» the company posted in June, crediting the model with «an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic.»

                                      The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it.

                                      There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of it indirectly. Nothing else on offer captures it at all.

                                      Open weights helped make FLUX an industry standard

                                      BFL officially launched in summer 2024 and gained a name for itself in the AI industry in the intervening two years for its commitment to open sourcing high-quality AI image models beloved by developers, creatives, and enterprises.

                                      The company’s founders, including Rombach, Andreas Blattmann and Patrick Esser, previously helped create VQGAN, latent diffusion and Stable Diffusion, the latter the open source technology that kicked off broad AI generation capabilities for the masses and currently used by many AI image generators and companies.

                                      That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart and Nous Research’s Hermes Agent, among other platforms, and the company cites film director Martin Scorsese among professional users.

                                      Wired magazine described Black Forest Labs as a relatively small company that nevertheless became a leading competitor to Silicon Valley’s largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on AI code sharing community Hugging Face. The company says it now runs a 100-person team across Freiburg and San Francisco.

                                      FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev and related control models, released shortly after the firm’s launch, gave researchers and creative-tool developers access to downloadable checkpoints, local inference and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license.

                                      The company continued that pattern with FLUX.2 Dev in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. Black Forest Labs called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code and optimized implementations for consumer Nvidia GPUs.

                                      FLUX 3 Dev raises the stakes on that evaluation. Previous Dev releases were image models. This one is described as a multimodal backbone spanning video, audio, image and action prediction — meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery. BFL hasn’t yet shared information about its license, the parameter count, quantizations or hardware requirements.

                                      The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and let teams adapt FLUX 3 to their own data, products and workflows.

                                      The financial backing behind FLUX 3 is worth noting alongside the technical claims. Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva and Deutsche Telekom’s T.Capital.

                                      ● Canal oficial · Gratis
                                      ¡Recibe las noticias antes que nadie!
                                      Únete a nuestro canal de WhatsApp y mantente informado al instante, sin spam.
                                      Unirme ahora →
                                      ● Noticias al instante ● Cobertura nacional ● Periodismo real Despertar Matinal
                                      — Redacción Despertar Matinal

                                      — Redacción Despertar Matinal

                                      Programa radial que te conecta con la información desde temprano en la mañana.

                                      Next Post
                                      Los posteos Anti-Argentina en X se viralizaron mediante casas de apuestas y bots

                                      Los posteos Anti-Argentina en X se viralizaron mediante casas de apuestas y bots

                                      Deja una respuesta Cancelar la respuesta

                                      Tu dirección de correo electrónico no será publicada. Los campos obligatorios están marcados con *

                                      Canal de WhatsApp

                                      WhatsApp logo WhatsApp

                                      Canal · Despertar Matinal

                                      Únete a nuestro
                                      Canal

                                      Seguir ahora

                                      El clima

                                      Canal de YouTube

                                      YouTube

                                      Canal · Despertar Matinal

                                      Mira nuestro
                                      Canal

                                      Ver ahora

                                      Escúchanos en Spotify

                                      Spotify

                                      Podcast · Despertar Matinal

                                      Escucha nuestro
                                      Podcast

                                      Escuchar ahora

                                      Noticias Populares

                                      • Harvard agrees to pay millions after morgue manager sold body parts

                                        Harvard agrees to pay millions after morgue manager sold body parts

                                        0 shares
                                        Share 0 Tweet 0
                                      • China recupera la etapa de un cohete en tierra por primera vez, después de una recuperación anterior en el mar

                                        0 shares
                                        Share 0 Tweet 0
                                      • Tesla se prepara para lanzar el Cybercab, pero crecen las dudas sobre si está listo

                                        0 shares
                                        Share 0 Tweet 0
                                      • Target receives $1bn boost from Trump tariff refunds

                                        0 shares
                                        Share 0 Tweet 0
                                      • ANSES hoy: quiénes cobran este miércoles 19 de agosto

                                        0 shares
                                        Share 0 Tweet 0

                                      Medio digital independiente con análisis, opinión y periodismo responsable desde República Dominicana.

                                      Secciones populares

                                      • Política
                                      • Economía & Negocios
                                      • Justicia
                                      • Turismo
                                      • Tecnología
                                      • Entretenimiento
                                      • Mundo
                                      • Cine y Series
                                      • Música
                                      • Moda

                                      Contenido

                                      • Titulares del Día
                                      • Mundo
                                      • Nacionales
                                      • Política
                                      • Deportes
                                      • Economía & Negocios
                                      • Ciencia
                                      • Entretenimiento
                                      • Podcast
                                      • Opinión
                                      • Despertar Matinal TV
                                      • Editoriales

                                      Corporativo

                                      • Sobre nosotros
                                      • Publicidad
                                      • Sala de prensa
                                      • Contacto
                                      • Política de Privacidad
                                      • Eliminación de Datos

                                      Boletines

                                      Suscríbete a nuestro boletín
                                      Recibe las noticias más importantes cada mañana.

                                      • Nosotros
                                      • Publicidad
                                      • Trabaja con nosotros
                                      • Contactos

                                      © 2025 Despertar Matinal. Aviso Legal - comunícate con nuestra redacción y obtén más información sobre Despertar Matinal..

                                      No Result
                                      View All Result
                                      • Home

                                      © 2025 Despertar Matinal. Aviso Legal - comunícate con nuestra redacción y obtén más información sobre Despertar Matinal..

                                      Welcome Back!

                                      Login to your account below

                                      Forgotten Password?

                                      Retrieve your password

                                      Please enter your username or email address to reset your password.

                                      Log In

                                      Desarrollado por
                                      ►
                                      Las cookies necesarias habilitan funciones esenciales del sitio como inicios de sesión seguros y ajustes de preferencias de consentimiento. No almacenan datos personales.
                                      Ninguno
                                      ►
                                      Las cookies funcionales soportan funciones como compartir contenido en redes sociales, recopilar comentarios y habilitar herramientas de terceros.
                                      Ninguno
                                      ►
                                      Las cookies analíticas rastrean las interacciones de los visitantes, proporcionando información sobre métricas como el número de visitantes, la tasa de rebote y las fuentes de tráfico.
                                      Ninguno
                                      ►
                                      Las cookies de publicidad ofrecen anuncios personalizados basados en tus visitas anteriores y analizan la efectividad de las campañas publicitarias.
                                      Ninguno
                                      ►
                                      Las cookies no clasificadas son aquellas que estamos en proceso de clasificar, junto con los proveedores de cookies individuales.
                                      Ninguno
                                      Desarrollado por