• Nosotros
  • Publicidad
  • Trabaja con nosotros
  • Contactos
domingo, agosto 16, 2026
  • Login
No Result
View All Result
NEWSLETTER
Despertar Matinal
  • Titulares del Día
    • All
    • En Portada
    PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

    PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

    Video- Presidente ADP advierte que mayoría de las escuelas públicas no tienen condiciones para resistir un terremoto

    Video- Presidente ADP advierte que mayoría de las escuelas públicas no tienen condiciones para resistir un terremoto

    Presidente Abinader pide agricultores tecnificarse para eliminar mano de obra extranjera

    Presidente Abinader pide agricultores tecnificarse para eliminar mano de obra extranjera

    Presidente Luis Abinader entrega polideportivo techado en el Centro Educativo Santo Cura de Ars

    Presidente Luis Abinader entrega polideportivo techado en el Centro Educativo Santo Cura de Ars

    Roberto Ángel Salcedo destaca avances de la cultura durante seis años de Gobierno de Abinader

    Roberto Ángel Salcedo destaca avances de la cultura durante seis años de Gobierno de Abinader

    Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

    Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

    Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

    Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

    Derechos Humanos y Comisión de la Verdad reiteran reclamo de justicia por explosión de San Cristóbal

    Derechos Humanos y Comisión de la Verdad reiteran reclamo de justicia por explosión de San Cristóbal

    ADOCALZA respalda mecanismos de INABIE y defiende capacidad de fabricantes nacionales

    ADOCALZA respalda mecanismos de INABIE y defiende capacidad de fabricantes nacionales

    Trending Tags

    • Mundo
      • All
      • América Latina
      • Conflictos Internacionales
      • Estados Unidos
      • Europa
      • Geopolítica
      • Haití
      • Medio Oriente
      Evangelina Anderson recordó el momento en que les contó a sus hijos sobre su separación: “No estaba preparada”

      Evangelina Anderson recordó el momento en que les contó a sus hijos sobre su separación: “No estaba preparada”

      Abelardo le ordenó por Twitter a una ministra echar a todos los izquierdistas de su ministerio

      Abelardo le ordenó por Twitter a una ministra echar a todos los izquierdistas de su ministerio

      La Tercera Ola del “socialismo cultural” en las universidades

      La Tercera Ola del “socialismo cultural” en las universidades

      La Guardia Civil confirmó 15 violaciones en Ceuta tras la invasión de inmigrantes ilegales marroquíes

      La Guardia Civil confirmó 15 violaciones en Ceuta tras la invasión de inmigrantes ilegales marroquíes

      El secretario de Asuntos Nucleares de Milei dejó en ridículo al kirchnerista Jorge Taiana

      El secretario de Asuntos Nucleares de Milei dejó en ridículo al kirchnerista Jorge Taiana

      La mentira noble nunca salva a quienes pretende proteger

      La mentira noble nunca salva a quienes pretende proteger

      Claude lanzará una marca de agua invisible para detectar textos generados con IA

      Claude lanzará una marca de agua invisible para detectar textos generados con IA

      San Martín, grande fue cuando el sol lo alumbraba y más grande en la puesta del sol

      San Martín, grande fue cuando el sol lo alumbraba y más grande en la puesta del sol

      Lula postergó la llegada del embajador de EEUU en Brasil hasta después de las elecciones

      Lula postergó la llegada del embajador de EEUU en Brasil hasta después de las elecciones

      Trending Tags

      • Nacionales
        • All
        • Bávaro Punta Cana
        • Educación
        • Gobierno
        • Infraestructura
        • Justicia
        • Obras Públicas
        • Opinión
        • Provincias
        • Seguridad Ciudadana
        • semana santa 2026
        • Sociedad
        • Transporte
        PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

        PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

        Video- Presidente ADP advierte que mayoría de las escuelas públicas no tienen condiciones para resistir un terremoto

        Video- Presidente ADP advierte que mayoría de las escuelas públicas no tienen condiciones para resistir un terremoto

        Presidente Abinader pide agricultores tecnificarse para eliminar mano de obra extranjera

        Presidente Abinader pide agricultores tecnificarse para eliminar mano de obra extranjera

        Presidente Luis Abinader entrega polideportivo techado en el Centro Educativo Santo Cura de Ars

        Presidente Luis Abinader entrega polideportivo techado en el Centro Educativo Santo Cura de Ars

        Entérese quienes se unen para impulsar el desarrollo...

        Entérese quienes se unen para impulsar el desarrollo…

        Roberto Ángel Salcedo destaca avances de la cultura durante seis años de Gobierno de Abinader

        Roberto Ángel Salcedo destaca avances de la cultura durante seis años de Gobierno de Abinader

        Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

        Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

        Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

        Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

        Intrant: nuevo sistema de licencias revierte pérdidas y genera...

        Intrant: nuevo sistema de licencias revierte pérdidas y genera…

        Trending Tags

        • Política
          • All
          • Congreso
          • Opinión Política
          • Partidos Políticos
          • Poder Municipal
          • Transparencia y Corrupción
          ARTICULO: De los millones de seguidores al poder: gobernar un país no es hacer un reality en YouTube

          ARTICULO: De los millones de seguidores al poder: gobernar un país no es hacer un reality en YouTube

          Estados Unidos no descarta operación militar contra Cuba

          Estados Unidos no descarta operación militar contra Cuba

          Tribunal Constitucional ratifica que País Posible es la 7ma fuerza...

          Tribunal Constitucional ratifica que País Posible es la 7ma fuerza…

          Sismo en Colombia suma 181 fallecidos

          Sismo en Colombia suma 181 fallecidos

          PLD dice Montecristi esta en el abandono; PRM promete obras

          PLD dice Montecristi esta en el abandono; PRM promete obras

          TSE rechaza suspender fondos públicos asignados a partidos en 2026

          TSE rechaza suspender fondos públicos asignados a partidos en 2026

          JCE impulsa debate regional sobre IA y transparencia electoral

          JCE impulsa debate regional sobre IA y transparencia electoral

          Reforma a Seguridad Social quedó fuera de agenda pese a promesa de Abinader – El Nuevo Diario (República Dominicana)

          Reforma a la seguridad social sigue sin llegar al Congreso

          ¡La dejaron pasar! Concluye otra legislatura sin aprobarse una reforma integral para erradicar los feminicidios en RD

          Congreso dominicano deja vencer, otra vez, la reforma urgente contra los feminicidios

          Trending Tags

          • Deportes
            • All
            • Atletas Dominicanos
            • Béisbol
            DR Open Kiteboarding Championship reúne atletas de 15 países y reafirma a Cabarete como capital del kitesurf del Caribe

            Cabarete se corona como capital histórica del kitesurf con el DR Open Championship 2026

            El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

            El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

            La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

            La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

            Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

            Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

            Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

            Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

            Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

            Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

            Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

            Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

            30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

            Buffalo recibe a Montreal para abrir la segunda ronda

            Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

            Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

            Trending Tags

            • Economía
              • All
              • Combustibles
              • Energía
              • Indicadores Económicos
              • Sector Energético
              • Turismo
              Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

              Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

              Aventúrate RD 2026

              Aventúrate RD 2026 revela agenda oficial y consolida el turismo de aventura dominicano

              WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

              WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

              Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

              Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

              Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

              Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

              Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

              Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

              El ministerio de Turismo cierra con éxito festival gastronómico “Saborea el Paraíso” en Sánchez, Samaná

              El Ministerio de Turismo celebra un exitoso cierre del festival gastronómico «Saborea el Paraíso» en Sánchez, Samaná

              El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

              El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

              Más allá del comercio: los efectos del arancel estadounidense sobre el turismo dominicano

              Arancel de EE.UU. pone a prueba al turismo dominicano y al silencio oficial del gobierno

              Trending Tags

              • Ciencia
                • All
                • Energía
                • Innovación
                • Investigación Científica
                • Salud y Medicina
                • Tecnología Médica
                Two dead and hundreds evacuated as wildfires break out near Athens

                Two dead and hundreds evacuated as wildfires break out near Athens

                PRD califica de “desastrosa” gestión de Abinader

                PRD califica de “desastrosa” gestión de Abinader

                Four 'extraordinary' Renaissance paintings stolen from Italian museum

                Four ‘extraordinary’ Renaissance paintings stolen from Italian museum

                Europe's tallest Virgin Mary statue unveiled in rural Poland

                Europe’s tallest Virgin Mary statue unveiled in rural Poland

                Three killed as Russia launches drone and missile attack on Ukraine

                Three killed as Russia launches drone and missile attack on Ukraine

                Twelve killed as Polish bus veers off Hungarian motorway

                Twelve killed as Polish bus veers off Hungarian motorway

                Storm Lala: Hawaii braces for potential first direct hit by a hurricane in 34 years

                Storm Lala: Hawaii braces for potential first direct hit by a hurricane in 34 years

                Rescuers search for survivors of powerful Indonesia earthquake

                Rescuers search for survivors of powerful Indonesia earthquake

                Australian state to begin gun buyback after Bondi Beach attack

                Australian state to begin gun buyback after Bondi Beach attack

                Trending Tags

                • Tecnología
                  • All
                  • Aplicaciones
                  • Inteligencia Artificial
                  El Flash V4 mejor clasificado de DeepSeek tropieza con tareas de agentes reales a medida que aumentan sus precios

                  El Flash V4 mejor clasificado de DeepSeek tropieza con tareas de agentes reales a medida que aumentan sus precios

                  Berkshire Hathaway aumentó su participación en Alphabet y constructoras de viviendas en el segundo trimestre

                  Berkshire Hathaway aumentó su participación en Alphabet y constructoras de viviendas en el segundo trimestre

                  Un equipo de evaluación encontró lo que la revisión cualitativa no pudo: los modelos de IA tienen más confianza cuando están equivocados

                  Un equipo de evaluación encontró lo que la revisión cualitativa no pudo: los modelos de IA tienen más confianza cuando están equivocados

                  El tribunal fiscal de Maryland anula el impuesto a la publicidad digital y ordena reembolsos a Apple, Google y Peacock TV

                  El tribunal fiscal de Maryland anula el impuesto a la publicidad digital y ordena reembolsos a Apple, Google y Peacock TV

                  GLM-5.3 está aquí con capacidades cibernéticas avanzadas y, según se informa, ya encontró una 'vulnerabilidad grave' en Cursor

                  GLM-5.3 está aquí con capacidades cibernéticas avanzadas y, según se informa, ya encontró una ‘vulnerabilidad grave’ en Cursor

                  El foro en línea Reddit se unirá al influyente y seguido de cerca índice S&P 500

                  El foro en línea Reddit se unirá al influyente y seguido de cerca índice S&P 500

                  30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                  ¿Un título de ‘influencer’? Universidades apuestan por especialización en creación de contenidos; Los críticos cuestionan el valor.

                  El terremoto de Colombia es un déjà vu para los venezolanos. La respuesta del gobierno es todo menos

                  El terremoto de Colombia es un déjà vu para los venezolanos. La respuesta del gobierno es todo menos

                  Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't tell users what they'd done

                  Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn’t tell users what they’d done

                  Trending Tags

                  • Entretenimiento
                    • All
                    • Cine y Series
                    • Cultura Digital
                    • Cultura Popular
                    • Gastronomía
                    • Música
                    Muere Bou Meng, artista camboyano y superviviente de un centro de tortura de los Jemeres Rojos, a los 85 años

                    Muere Bou Meng, artista camboyano y superviviente de un centro de tortura de los Jemeres Rojos, a los 85 años

                    Los fanáticos de Bonnie Tyler se alinean en las calles de un pueblo galés mientras traen su ataúd a casa.

                    Los fanáticos de Bonnie Tyler se alinean en las calles de un pueblo galés mientras traen su ataúd a casa.

                    Liechtenstein cambia sus reglas para permitir que las mujeres hereden el trono del principado alpino

                    Liechtenstein cambia sus reglas para permitir que las mujeres hereden el trono del principado alpino

                    Muere Mark Rydell, el director nominado al Oscar por 'En el estanque dorado', a los 97 años

                    Muere Mark Rydell, el director nominado al Oscar por ‘En el estanque dorado’, a los 97 años

                    Ellen Greene vuelve a visitar a Audrey de 'La pequeña tienda de los horrores' para el 40 aniversario de la película

                    Ellen Greene vuelve a visitar a Audrey de ‘La pequeña tienda de los horrores’ para el 40 aniversario de la película

                    La Sra. Lauryn Hill y el compañero de banda de Fugees, Wyclef Jean, encabezarán el Global Citizen Festival

                    La Sra. Lauryn Hill y el compañero de banda de Fugees, Wyclef Jean, encabezarán el Global Citizen Festival

                    Marc Anthony, Chayanne y más darán un concierto benéfico para ayudar en el terremoto de Venezuela y Colombia

                    Marc Anthony, Chayanne y más darán un concierto benéfico para ayudar en el terremoto de Venezuela y Colombia

                    Reseña musical: 'Comes in Waves' de Carly Simon es un viaje a través del amor y la pérdida

                    Reseña musical: ‘Comes in Waves’ de Carly Simon es un viaje a través del amor y la pérdida

                    Reseña de la película: 'El fin de Oak Street' es un buen momento gonzo

                    Reseña de la película: ‘El fin de Oak Street’ es un buen momento gonzo

                    Trending Tags

                    • Titulares del Día
                      • All
                      • En Portada
                      PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

                      PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

                      Video- Presidente ADP advierte que mayoría de las escuelas públicas no tienen condiciones para resistir un terremoto

                      Video- Presidente ADP advierte que mayoría de las escuelas públicas no tienen condiciones para resistir un terremoto

                      Presidente Abinader pide agricultores tecnificarse para eliminar mano de obra extranjera

                      Presidente Abinader pide agricultores tecnificarse para eliminar mano de obra extranjera

                      Presidente Luis Abinader entrega polideportivo techado en el Centro Educativo Santo Cura de Ars

                      Presidente Luis Abinader entrega polideportivo techado en el Centro Educativo Santo Cura de Ars

                      Roberto Ángel Salcedo destaca avances de la cultura durante seis años de Gobierno de Abinader

                      Roberto Ángel Salcedo destaca avances de la cultura durante seis años de Gobierno de Abinader

                      Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

                      Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

                      Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

                      Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

                      Derechos Humanos y Comisión de la Verdad reiteran reclamo de justicia por explosión de San Cristóbal

                      Derechos Humanos y Comisión de la Verdad reiteran reclamo de justicia por explosión de San Cristóbal

                      ADOCALZA respalda mecanismos de INABIE y defiende capacidad de fabricantes nacionales

                      ADOCALZA respalda mecanismos de INABIE y defiende capacidad de fabricantes nacionales

                      Trending Tags

                      • Mundo
                        • All
                        • América Latina
                        • Conflictos Internacionales
                        • Estados Unidos
                        • Europa
                        • Geopolítica
                        • Haití
                        • Medio Oriente
                        Evangelina Anderson recordó el momento en que les contó a sus hijos sobre su separación: “No estaba preparada”

                        Evangelina Anderson recordó el momento en que les contó a sus hijos sobre su separación: “No estaba preparada”

                        Abelardo le ordenó por Twitter a una ministra echar a todos los izquierdistas de su ministerio

                        Abelardo le ordenó por Twitter a una ministra echar a todos los izquierdistas de su ministerio

                        La Tercera Ola del “socialismo cultural” en las universidades

                        La Tercera Ola del “socialismo cultural” en las universidades

                        La Guardia Civil confirmó 15 violaciones en Ceuta tras la invasión de inmigrantes ilegales marroquíes

                        La Guardia Civil confirmó 15 violaciones en Ceuta tras la invasión de inmigrantes ilegales marroquíes

                        El secretario de Asuntos Nucleares de Milei dejó en ridículo al kirchnerista Jorge Taiana

                        El secretario de Asuntos Nucleares de Milei dejó en ridículo al kirchnerista Jorge Taiana

                        La mentira noble nunca salva a quienes pretende proteger

                        La mentira noble nunca salva a quienes pretende proteger

                        Claude lanzará una marca de agua invisible para detectar textos generados con IA

                        Claude lanzará una marca de agua invisible para detectar textos generados con IA

                        San Martín, grande fue cuando el sol lo alumbraba y más grande en la puesta del sol

                        San Martín, grande fue cuando el sol lo alumbraba y más grande en la puesta del sol

                        Lula postergó la llegada del embajador de EEUU en Brasil hasta después de las elecciones

                        Lula postergó la llegada del embajador de EEUU en Brasil hasta después de las elecciones

                        Trending Tags

                        • Nacionales
                          • All
                          • Bávaro Punta Cana
                          • Educación
                          • Gobierno
                          • Infraestructura
                          • Justicia
                          • Obras Públicas
                          • Opinión
                          • Provincias
                          • Seguridad Ciudadana
                          • semana santa 2026
                          • Sociedad
                          • Transporte
                          PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

                          PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

                          Video- Presidente ADP advierte que mayoría de las escuelas públicas no tienen condiciones para resistir un terremoto

                          Video- Presidente ADP advierte que mayoría de las escuelas públicas no tienen condiciones para resistir un terremoto

                          Presidente Abinader pide agricultores tecnificarse para eliminar mano de obra extranjera

                          Presidente Abinader pide agricultores tecnificarse para eliminar mano de obra extranjera

                          Presidente Luis Abinader entrega polideportivo techado en el Centro Educativo Santo Cura de Ars

                          Presidente Luis Abinader entrega polideportivo techado en el Centro Educativo Santo Cura de Ars

                          Entérese quienes se unen para impulsar el desarrollo...

                          Entérese quienes se unen para impulsar el desarrollo…

                          Roberto Ángel Salcedo destaca avances de la cultura durante seis años de Gobierno de Abinader

                          Roberto Ángel Salcedo destaca avances de la cultura durante seis años de Gobierno de Abinader

                          Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

                          Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

                          Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

                          Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

                          Intrant: nuevo sistema de licencias revierte pérdidas y genera...

                          Intrant: nuevo sistema de licencias revierte pérdidas y genera…

                          Trending Tags

                          • Política
                            • All
                            • Congreso
                            • Opinión Política
                            • Partidos Políticos
                            • Poder Municipal
                            • Transparencia y Corrupción
                            ARTICULO: De los millones de seguidores al poder: gobernar un país no es hacer un reality en YouTube

                            ARTICULO: De los millones de seguidores al poder: gobernar un país no es hacer un reality en YouTube

                            Estados Unidos no descarta operación militar contra Cuba

                            Estados Unidos no descarta operación militar contra Cuba

                            Tribunal Constitucional ratifica que País Posible es la 7ma fuerza...

                            Tribunal Constitucional ratifica que País Posible es la 7ma fuerza…

                            Sismo en Colombia suma 181 fallecidos

                            Sismo en Colombia suma 181 fallecidos

                            PLD dice Montecristi esta en el abandono; PRM promete obras

                            PLD dice Montecristi esta en el abandono; PRM promete obras

                            TSE rechaza suspender fondos públicos asignados a partidos en 2026

                            TSE rechaza suspender fondos públicos asignados a partidos en 2026

                            JCE impulsa debate regional sobre IA y transparencia electoral

                            JCE impulsa debate regional sobre IA y transparencia electoral

                            Reforma a Seguridad Social quedó fuera de agenda pese a promesa de Abinader – El Nuevo Diario (República Dominicana)

                            Reforma a la seguridad social sigue sin llegar al Congreso

                            ¡La dejaron pasar! Concluye otra legislatura sin aprobarse una reforma integral para erradicar los feminicidios en RD

                            Congreso dominicano deja vencer, otra vez, la reforma urgente contra los feminicidios

                            Trending Tags

                            • Deportes
                              • All
                              • Atletas Dominicanos
                              • Béisbol
                              DR Open Kiteboarding Championship reúne atletas de 15 países y reafirma a Cabarete como capital del kitesurf del Caribe

                              Cabarete se corona como capital histórica del kitesurf con el DR Open Championship 2026

                              El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

                              El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

                              La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

                              La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

                              Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

                              Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

                              Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

                              Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

                              Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

                              Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

                              Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

                              Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

                              30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                              Buffalo recibe a Montreal para abrir la segunda ronda

                              Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

                              Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

                              Trending Tags

                              • Economía
                                • All
                                • Combustibles
                                • Energía
                                • Indicadores Económicos
                                • Sector Energético
                                • Turismo
                                Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

                                Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

                                Aventúrate RD 2026

                                Aventúrate RD 2026 revela agenda oficial y consolida el turismo de aventura dominicano

                                WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

                                WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

                                Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

                                Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

                                Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

                                Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

                                Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

                                Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

                                El ministerio de Turismo cierra con éxito festival gastronómico “Saborea el Paraíso” en Sánchez, Samaná

                                El Ministerio de Turismo celebra un exitoso cierre del festival gastronómico «Saborea el Paraíso» en Sánchez, Samaná

                                El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

                                El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

                                Más allá del comercio: los efectos del arancel estadounidense sobre el turismo dominicano

                                Arancel de EE.UU. pone a prueba al turismo dominicano y al silencio oficial del gobierno

                                Trending Tags

                                • Ciencia
                                  • All
                                  • Energía
                                  • Innovación
                                  • Investigación Científica
                                  • Salud y Medicina
                                  • Tecnología Médica
                                  Two dead and hundreds evacuated as wildfires break out near Athens

                                  Two dead and hundreds evacuated as wildfires break out near Athens

                                  PRD califica de “desastrosa” gestión de Abinader

                                  PRD califica de “desastrosa” gestión de Abinader

                                  Four 'extraordinary' Renaissance paintings stolen from Italian museum

                                  Four ‘extraordinary’ Renaissance paintings stolen from Italian museum

                                  Europe's tallest Virgin Mary statue unveiled in rural Poland

                                  Europe’s tallest Virgin Mary statue unveiled in rural Poland

                                  Three killed as Russia launches drone and missile attack on Ukraine

                                  Three killed as Russia launches drone and missile attack on Ukraine

                                  Twelve killed as Polish bus veers off Hungarian motorway

                                  Twelve killed as Polish bus veers off Hungarian motorway

                                  Storm Lala: Hawaii braces for potential first direct hit by a hurricane in 34 years

                                  Storm Lala: Hawaii braces for potential first direct hit by a hurricane in 34 years

                                  Rescuers search for survivors of powerful Indonesia earthquake

                                  Rescuers search for survivors of powerful Indonesia earthquake

                                  Australian state to begin gun buyback after Bondi Beach attack

                                  Australian state to begin gun buyback after Bondi Beach attack

                                  Trending Tags

                                  • Tecnología
                                    • All
                                    • Aplicaciones
                                    • Inteligencia Artificial
                                    El Flash V4 mejor clasificado de DeepSeek tropieza con tareas de agentes reales a medida que aumentan sus precios

                                    El Flash V4 mejor clasificado de DeepSeek tropieza con tareas de agentes reales a medida que aumentan sus precios

                                    Berkshire Hathaway aumentó su participación en Alphabet y constructoras de viviendas en el segundo trimestre

                                    Berkshire Hathaway aumentó su participación en Alphabet y constructoras de viviendas en el segundo trimestre

                                    Un equipo de evaluación encontró lo que la revisión cualitativa no pudo: los modelos de IA tienen más confianza cuando están equivocados

                                    Un equipo de evaluación encontró lo que la revisión cualitativa no pudo: los modelos de IA tienen más confianza cuando están equivocados

                                    El tribunal fiscal de Maryland anula el impuesto a la publicidad digital y ordena reembolsos a Apple, Google y Peacock TV

                                    El tribunal fiscal de Maryland anula el impuesto a la publicidad digital y ordena reembolsos a Apple, Google y Peacock TV

                                    GLM-5.3 está aquí con capacidades cibernéticas avanzadas y, según se informa, ya encontró una 'vulnerabilidad grave' en Cursor

                                    GLM-5.3 está aquí con capacidades cibernéticas avanzadas y, según se informa, ya encontró una ‘vulnerabilidad grave’ en Cursor

                                    El foro en línea Reddit se unirá al influyente y seguido de cerca índice S&P 500

                                    El foro en línea Reddit se unirá al influyente y seguido de cerca índice S&P 500

                                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                    ¿Un título de ‘influencer’? Universidades apuestan por especialización en creación de contenidos; Los críticos cuestionan el valor.

                                    El terremoto de Colombia es un déjà vu para los venezolanos. La respuesta del gobierno es todo menos

                                    El terremoto de Colombia es un déjà vu para los venezolanos. La respuesta del gobierno es todo menos

                                    Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't tell users what they'd done

                                    Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn’t tell users what they’d done

                                    Trending Tags

                                    • Entretenimiento
                                      • All
                                      • Cine y Series
                                      • Cultura Digital
                                      • Cultura Popular
                                      • Gastronomía
                                      • Música
                                      Muere Bou Meng, artista camboyano y superviviente de un centro de tortura de los Jemeres Rojos, a los 85 años

                                      Muere Bou Meng, artista camboyano y superviviente de un centro de tortura de los Jemeres Rojos, a los 85 años

                                      Los fanáticos de Bonnie Tyler se alinean en las calles de un pueblo galés mientras traen su ataúd a casa.

                                      Los fanáticos de Bonnie Tyler se alinean en las calles de un pueblo galés mientras traen su ataúd a casa.

                                      Liechtenstein cambia sus reglas para permitir que las mujeres hereden el trono del principado alpino

                                      Liechtenstein cambia sus reglas para permitir que las mujeres hereden el trono del principado alpino

                                      Muere Mark Rydell, el director nominado al Oscar por 'En el estanque dorado', a los 97 años

                                      Muere Mark Rydell, el director nominado al Oscar por ‘En el estanque dorado’, a los 97 años

                                      Ellen Greene vuelve a visitar a Audrey de 'La pequeña tienda de los horrores' para el 40 aniversario de la película

                                      Ellen Greene vuelve a visitar a Audrey de ‘La pequeña tienda de los horrores’ para el 40 aniversario de la película

                                      La Sra. Lauryn Hill y el compañero de banda de Fugees, Wyclef Jean, encabezarán el Global Citizen Festival

                                      La Sra. Lauryn Hill y el compañero de banda de Fugees, Wyclef Jean, encabezarán el Global Citizen Festival

                                      Marc Anthony, Chayanne y más darán un concierto benéfico para ayudar en el terremoto de Venezuela y Colombia

                                      Marc Anthony, Chayanne y más darán un concierto benéfico para ayudar en el terremoto de Venezuela y Colombia

                                      Reseña musical: 'Comes in Waves' de Carly Simon es un viaje a través del amor y la pérdida

                                      Reseña musical: ‘Comes in Waves’ de Carly Simon es un viaje a través del amor y la pérdida

                                      Reseña de la película: 'El fin de Oak Street' es un buen momento gonzo

                                      Reseña de la película: ‘El fin de Oak Street’ es un buen momento gonzo

                                      Trending Tags

                                      No Result
                                      View All Result
                                      Despertar Matinal
                                      No Result
                                      View All Result

                                      Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn’t tell users what they’d done

                                      by — Redacción Despertar Matinal
                                      13 de agosto de 2026
                                      in Tecnología
                                      0
                                      Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't tell users what they'd done
                                      0
                                      SHARES
                                      1
                                      VIEWS
                                      Share on FacebookShare on Twitter

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      Tours Colombia Todo el año Tours Colombia Todo el año Tours Colombia Todo el año

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      ¡No te pierdas las noticias destacadas!

                                      Suscríbete y recibe las historias más importantes del día.

                                      Al suscribirte aceptas nuestros términos y condiciones y política de privacidad.

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      Tours Colombia Todo el año Tours Colombia Todo el año Tours Colombia Todo el año

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      ● Canal oficial · Gratis
                                      ¡Recibe las noticias antes que nadie!
                                      Únete a nuestro canal de WhatsApp y mantente informado al instante, sin spam.
                                      Unirme ahora →
                                      ● Noticias al instante ● Cobertura nacional ● Periodismo real Despertar Matinal
                                      — Redacción Despertar Matinal

                                      — Redacción Despertar Matinal

                                      Programa radial que te conecta con la información desde temprano en la mañana.

                                      Next Post
                                      Johnny Pujols: el PLD no necesita mentir para cuestionar al Gobierno

                                      Johnny Pujols: el PLD no necesita mentir para cuestionar al Gobierno

                                      Deja una respuesta Cancelar la respuesta

                                      Tu dirección de correo electrónico no será publicada. Los campos obligatorios están marcados con *

                                      Canal de WhatsApp

                                      WhatsApp logo WhatsApp

                                      Canal · Despertar Matinal

                                      Únete a nuestro
                                      Canal

                                      Seguir ahora

                                      El clima

                                      Canal de YouTube

                                      YouTube

                                      Canal · Despertar Matinal

                                      Mira nuestro
                                      Canal

                                      Ver ahora

                                      Escúchanos en Spotify

                                      Spotify

                                      Podcast · Despertar Matinal

                                      Escucha nuestro
                                      Podcast

                                      Escuchar ahora

                                      Noticias Populares

                                      • PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

                                        PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

                                        0 shares
                                        Share 0 Tweet 0
                                      • Twelve killed as Polish bus veers off Hungarian motorway

                                        0 shares
                                        Share 0 Tweet 0
                                      • Cristiano Ronaldo marries long-time partner Georgina Rodríguez

                                        0 shares
                                        Share 0 Tweet 0
                                      • Donald Trump declarará al estrecho de Ormuz como territorio de los Estados Unidos

                                        0 shares
                                        Share 0 Tweet 0
                                      • Claude lanzará una marca de agua invisible para detectar textos generados con IA

                                        0 shares
                                        Share 0 Tweet 0

                                      Medio digital independiente con análisis, opinión y periodismo responsable desde República Dominicana.

                                      Secciones populares

                                      • Política
                                      • Economía & Negocios
                                      • Justicia
                                      • Turismo
                                      • Tecnología
                                      • Entretenimiento
                                      • Mundo
                                      • Cine y Series
                                      • Música
                                      • Moda

                                      Contenido

                                      • Titulares del Día
                                      • Mundo
                                      • Nacionales
                                      • Política
                                      • Deportes
                                      • Economía & Negocios
                                      • Ciencia
                                      • Entretenimiento
                                      • Podcast
                                      • Opinión
                                      • Despertar Matinal TV
                                      • Editoriales

                                      Corporativo

                                      • Sobre nosotros
                                      • Publicidad
                                      • Sala de prensa
                                      • Contacto
                                      • Política de Privacidad
                                      • Eliminación de Datos

                                      Boletines

                                      Suscríbete a nuestro boletín
                                      Recibe las noticias más importantes cada mañana.

                                      • Nosotros
                                      • Publicidad
                                      • Trabaja con nosotros
                                      • Contactos

                                      © 2025 Despertar Matinal. Aviso Legal - comunícate con nuestra redacción y obtén más información sobre Despertar Matinal..

                                      No Result
                                      View All Result
                                      • Home

                                      © 2025 Despertar Matinal. Aviso Legal - comunícate con nuestra redacción y obtén más información sobre Despertar Matinal..

                                      Welcome Back!

                                      Login to your account below

                                      Forgotten Password?

                                      Retrieve your password

                                      Please enter your username or email address to reset your password.

                                      Log In

                                      Desarrollado por
                                      ►
                                      Las cookies necesarias habilitan funciones esenciales del sitio como inicios de sesión seguros y ajustes de preferencias de consentimiento. No almacenan datos personales.
                                      Ninguno
                                      ►
                                      Las cookies funcionales soportan funciones como compartir contenido en redes sociales, recopilar comentarios y habilitar herramientas de terceros.
                                      Ninguno
                                      ►
                                      Las cookies analíticas rastrean las interacciones de los visitantes, proporcionando información sobre métricas como el número de visitantes, la tasa de rebote y las fuentes de tráfico.
                                      Ninguno
                                      ►
                                      Las cookies de publicidad ofrecen anuncios personalizados basados en tus visitas anteriores y analizan la efectividad de las campañas publicitarias.
                                      Ninguno
                                      ►
                                      Las cookies no clasificadas son aquellas que estamos en proceso de clasificar, junto con los proveedores de cookies individuales.
                                      Ninguno
                                      Desarrollado por