• Nosotros
  • Publicidad
  • Trabaja con nosotros
  • Contactos
sábado, agosto 15, 2026
  • Login
No Result
View All Result
NEWSLETTER
Despertar Matinal
  • Titulares del Día
    • All
    • En Portada
    Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

    Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

    Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

    Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

    Derechos Humanos y Comisión de la Verdad reiteran reclamo de justicia por explosión de San Cristóbal

    Derechos Humanos y Comisión de la Verdad reiteran reclamo de justicia por explosión de San Cristóbal

    ADOCALZA respalda mecanismos de INABIE y defiende capacidad de fabricantes nacionales

    ADOCALZA respalda mecanismos de INABIE y defiende capacidad de fabricantes nacionales

    Diputado Wendy Batista: “Código Penal lo satanizaron; Cree en acumulación de penas, en algunos casos

    Diputado Wendy Batista: “Código Penal lo satanizaron; Cree en acumulación de penas, en algunos casos

    “El PLD es el único partido que ha crecido y va a ganar en 2028”, afirma Johnny Pujols

    Johnny Pujols afirma PLD está preparada para disputar el poder al oficialismo en 2028

    Onesvie redobla prevención ante terremotos y suma más de 4,100 evaluaciones

    Onesvie redobla prevención ante terremotos y suma más de 4,100 evaluaciones

    ADOCCO rechaza propuesta de entregar a un civil el manejo financiero y de compras de la Policía Nacional

    ADOCCO rechaza propuesta de entregar a un civil el manejo financiero y de compras de la Policía Nacional

    Quique Antún advierte que República Dominicana no está preparada para un terremoto de gran magnitud

    Quique Antún advierte que República Dominicana no está preparada para un terremoto de gran magnitud

    Trending Tags

    • Mundo
      • All
      • América Latina
      • Conflictos Internacionales
      • Estados Unidos
      • Europa
      • Geopolítica
      • Haití
      • Medio Oriente
      Los Increíbles 3: cuándo se estrena y todos los detalles de la película más esperada de Pixar

      Los Increíbles 3: cuándo se estrena y todos los detalles de la película más esperada de Pixar

      El gobierno comunista de Pedro Sánchez refuerza la presencia militar en Ceuta ante amenazas de una nueva invasión

      El gobierno comunista de Pedro Sánchez refuerza la presencia militar en Ceuta ante amenazas de una nueva invasión

      Desastre: Casación ordenó que la causa de los supuestos testaferros de la AFA vuelva a Campana “sin dilaciones”

      Desastre: Casación ordenó que la causa de los supuestos testaferros de la AFA vuelva a Campana “sin dilaciones”

      Director de escuela rural en Misiones se tomó 18 años enteros de licencia y lo denuncian por defraudación

      Director de escuela rural en Misiones se tomó 18 años enteros de licencia y lo denuncian por defraudación

      Delcy Rodríguez anunció que hoy Venezuela liberará a 131 presos políticos

      Delcy Rodríguez anunció que hoy Venezuela liberará a 131 presos políticos

      Con un nuevo Centro Comercial, La Calera supera las 50 empresas instaladas desde el Régimen de Inversiones

      Con un nuevo Centro Comercial, La Calera supera las 50 empresas instaladas desde el Régimen de Inversiones

      Donald Trump declarará al estrecho de Ormuz como territorio de los Estados Unidos

      Donald Trump declarará al estrecho de Ormuz como territorio de los Estados Unidos

      PS Plus: todos los juegos que llegan en agosto y septiembre de 2026

      PS Plus: todos los juegos que llegan en agosto y septiembre de 2026

      Monotributo: qué pasa si ARCA te recategorizó de oficio

      Monotributo: qué pasa si ARCA te recategorizó de oficio

      Trending Tags

      • Nacionales
        • All
        • Bávaro Punta Cana
        • Educación
        • Gobierno
        • Infraestructura
        • Justicia
        • Obras Públicas
        • Opinión
        • Provincias
        • Seguridad Ciudadana
        • semana santa 2026
        • Sociedad
        • Transporte
        Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

        Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

        Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

        Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

        Intrant: nuevo sistema de licencias revierte pérdidas y genera...

        Intrant: nuevo sistema de licencias revierte pérdidas y genera…

        Derechos Humanos y Comisión de la Verdad reiteran reclamo de justicia por explosión de San Cristóbal

        Derechos Humanos y Comisión de la Verdad reiteran reclamo de justicia por explosión de San Cristóbal

        ADOCALZA respalda mecanismos de INABIE y defiende capacidad de fabricantes nacionales

        ADOCALZA respalda mecanismos de INABIE y defiende capacidad de fabricantes nacionales

        Tribunal levanta impedimento de salida a Jean Alain Rodríguez para estudios médicos

        Defensa de Jean Alain denuncia campaña para distraer debate…

        Avanza recuperación de terrenos de dominio público en Verón-Punta Cana

        Gobierno recupera terrenos públicos en Verón-Punta Cana tras años de ocupación

        Indotel presenta propuesta de nueva Ley de Telecomunicaciones...

        Indotel presenta propuesta de nueva Ley de Telecomunicaciones…

        Diputado Wendy Batista: “Código Penal lo satanizaron; Cree en acumulación de penas, en algunos casos

        Diputado Wendy Batista: “Código Penal lo satanizaron; Cree en acumulación de penas, en algunos casos

        Trending Tags

        • Política
          • All
          • Congreso
          • Opinión Política
          • Partidos Políticos
          • Poder Municipal
          • Transparencia y Corrupción
          Estados Unidos no descarta operación militar contra Cuba

          Estados Unidos no descarta operación militar contra Cuba

          Tribunal Constitucional ratifica que País Posible es la 7ma fuerza...

          Tribunal Constitucional ratifica que País Posible es la 7ma fuerza…

          Sismo en Colombia suma 181 fallecidos

          Sismo en Colombia suma 181 fallecidos

          PLD dice Montecristi esta en el abandono; PRM promete obras

          PLD dice Montecristi esta en el abandono; PRM promete obras

          TSE rechaza suspender fondos públicos asignados a partidos en 2026

          TSE rechaza suspender fondos públicos asignados a partidos en 2026

          JCE impulsa debate regional sobre IA y transparencia electoral

          JCE impulsa debate regional sobre IA y transparencia electoral

          Reforma a Seguridad Social quedó fuera de agenda pese a promesa de Abinader – El Nuevo Diario (República Dominicana)

          Reforma a la seguridad social sigue sin llegar al Congreso

          ¡La dejaron pasar! Concluye otra legislatura sin aprobarse una reforma integral para erradicar los feminicidios en RD

          Congreso dominicano deja vencer, otra vez, la reforma urgente contra los feminicidios

          Fuerza del Pueblo participa en Gran Desfile Dominicano del Bronx; llevó amplio contingente de simpatizantes

          Fuerza del Pueblo impresiona con masiva participación en el Desfile Dominicano del Bronx

          Trending Tags

          • Deportes
            • All
            • Atletas Dominicanos
            • Béisbol
            DR Open Kiteboarding Championship reúne atletas de 15 países y reafirma a Cabarete como capital del kitesurf del Caribe

            Cabarete se corona como capital histórica del kitesurf con el DR Open Championship 2026

            El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

            El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

            La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

            La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

            Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

            Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

            Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

            Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

            Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

            Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

            Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

            Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

            30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

            Buffalo recibe a Montreal para abrir la segunda ronda

            Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

            Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

            Trending Tags

            • Economía
              • All
              • Combustibles
              • Energía
              • Indicadores Económicos
              • Sector Energético
              • Turismo
              Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

              Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

              Aventúrate RD 2026

              Aventúrate RD 2026 revela agenda oficial y consolida el turismo de aventura dominicano

              WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

              WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

              Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

              Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

              Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

              Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

              Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

              Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

              El ministerio de Turismo cierra con éxito festival gastronómico “Saborea el Paraíso” en Sánchez, Samaná

              El Ministerio de Turismo celebra un exitoso cierre del festival gastronómico «Saborea el Paraíso» en Sánchez, Samaná

              El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

              El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

              Más allá del comercio: los efectos del arancel estadounidense sobre el turismo dominicano

              Arancel de EE.UU. pone a prueba al turismo dominicano y al silencio oficial del gobierno

              Trending Tags

              • Ciencia
                • All
                • Energía
                • Innovación
                • Investigación Científica
                • Salud y Medicina
                • Tecnología Médica
                Seven killed in Israeli strike on southern Lebanon, country's PM says

                Seven killed in Israeli strike on southern Lebanon, country’s PM says

                Family stranded at sea for 16 hours after jet ski capsized in Thailand

                Family stranded at sea for 16 hours after jet ski capsized in Thailand

                South Korea proposes talks to officially end war with North

                South Korea proposes talks to officially end war with North

                Five dead after 7.7 magnitude Indonesia earthquake

                Five dead after 7.7 magnitude Indonesia earthquake

                BBC seeks to subpoena Trump's family members in Panorama lawsuit

                BBC seeks to subpoena Trump’s family members in Panorama lawsuit

                The Polygamist's creator says women see themselves reflected in her Netflix hit

                The Polygamist’s creator says women see themselves reflected in her Netflix hit

                Dozens injured and thousands evacuated in Croatia wildfire

                Dozens injured and thousands evacuated in Croatia wildfire

                Luigi Mangione admits killing healthcare CEO and pleads guilty to federal charges

                Luigi Mangione admits killing healthcare CEO and pleads guilty to federal charges

                Afghan women tell the BBC their lives are unrecognisable after five years of Taliban rule

                Afghan women tell the BBC their lives are unrecognisable after five years of Taliban rule

                Trending Tags

                • Tecnología
                  • All
                  • Aplicaciones
                  • Inteligencia Artificial
                  El tribunal fiscal de Maryland anula el impuesto a la publicidad digital y ordena reembolsos a Apple, Google y Peacock TV

                  El tribunal fiscal de Maryland anula el impuesto a la publicidad digital y ordena reembolsos a Apple, Google y Peacock TV

                  GLM-5.3 está aquí con capacidades cibernéticas avanzadas y, según se informa, ya encontró una 'vulnerabilidad grave' en Cursor

                  GLM-5.3 está aquí con capacidades cibernéticas avanzadas y, según se informa, ya encontró una ‘vulnerabilidad grave’ en Cursor

                  El foro en línea Reddit se unirá al influyente y seguido de cerca índice S&P 500

                  El foro en línea Reddit se unirá al influyente y seguido de cerca índice S&P 500

                  30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                  ¿Un título de ‘influencer’? Universidades apuestan por especialización en creación de contenidos; Los críticos cuestionan el valor.

                  El terremoto de Colombia es un déjà vu para los venezolanos. La respuesta del gobierno es todo menos

                  El terremoto de Colombia es un déjà vu para los venezolanos. La respuesta del gobierno es todo menos

                  Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't tell users what they'd done

                  Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn’t tell users what they’d done

                  Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut

                  Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut

                  DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices

                  DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices

                  Por qué Capital One construyó su plataforma de IA multiagente en torno a modelos abiertos

                  Por qué Capital One construyó su plataforma de IA multiagente en torno a modelos abiertos

                  Trending Tags

                  • Entretenimiento
                    • All
                    • Cine y Series
                    • Cultura Digital
                    • Cultura Popular
                    • Gastronomía
                    • Música
                    Muere Mark Rydell, el director nominado al Oscar por 'En el estanque dorado', a los 97 años

                    Muere Mark Rydell, el director nominado al Oscar por ‘En el estanque dorado’, a los 97 años

                    Ellen Greene vuelve a visitar a Audrey de 'La pequeña tienda de los horrores' para el 40 aniversario de la película

                    Ellen Greene vuelve a visitar a Audrey de ‘La pequeña tienda de los horrores’ para el 40 aniversario de la película

                    La Sra. Lauryn Hill y el compañero de banda de Fugees, Wyclef Jean, encabezarán el Global Citizen Festival

                    La Sra. Lauryn Hill y el compañero de banda de Fugees, Wyclef Jean, encabezarán el Global Citizen Festival

                    Marc Anthony, Chayanne y más darán un concierto benéfico para ayudar en el terremoto de Venezuela y Colombia

                    Marc Anthony, Chayanne y más darán un concierto benéfico para ayudar en el terremoto de Venezuela y Colombia

                    Reseña musical: 'Comes in Waves' de Carly Simon es un viaje a través del amor y la pérdida

                    Reseña musical: ‘Comes in Waves’ de Carly Simon es un viaje a través del amor y la pérdida

                    Reseña de la película: 'El fin de Oak Street' es un buen momento gonzo

                    Reseña de la película: ‘El fin de Oak Street’ es un buen momento gonzo

                    Sir Anthony Hopkins sobre su álbum debut: 'Mi primer amor es la música'

                    Sir Anthony Hopkins sobre su álbum debut: ‘Mi primer amor es la música’

                    Taylor Swift, Lyle Lovett y más se dirigen al Salón de la Fama de los Compositores de Nashville

                    Taylor Swift, Lyle Lovett y más se dirigen al Salón de la Fama de los Compositores de Nashville

                    'The Galloping Cure', la ópera de Missy Mazzoli sobre la epidemia de opioides, se estrena en el Festival de Edimburgo

                    ‘The Galloping Cure’, la ópera de Missy Mazzoli sobre la epidemia de opioides, se estrena en el Festival de Edimburgo

                    Trending Tags

                    • Titulares del Día
                      • All
                      • En Portada
                      Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

                      Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

                      Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

                      Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

                      Derechos Humanos y Comisión de la Verdad reiteran reclamo de justicia por explosión de San Cristóbal

                      Derechos Humanos y Comisión de la Verdad reiteran reclamo de justicia por explosión de San Cristóbal

                      ADOCALZA respalda mecanismos de INABIE y defiende capacidad de fabricantes nacionales

                      ADOCALZA respalda mecanismos de INABIE y defiende capacidad de fabricantes nacionales

                      Diputado Wendy Batista: “Código Penal lo satanizaron; Cree en acumulación de penas, en algunos casos

                      Diputado Wendy Batista: “Código Penal lo satanizaron; Cree en acumulación de penas, en algunos casos

                      “El PLD es el único partido que ha crecido y va a ganar en 2028”, afirma Johnny Pujols

                      Johnny Pujols afirma PLD está preparada para disputar el poder al oficialismo en 2028

                      Onesvie redobla prevención ante terremotos y suma más de 4,100 evaluaciones

                      Onesvie redobla prevención ante terremotos y suma más de 4,100 evaluaciones

                      ADOCCO rechaza propuesta de entregar a un civil el manejo financiero y de compras de la Policía Nacional

                      ADOCCO rechaza propuesta de entregar a un civil el manejo financiero y de compras de la Policía Nacional

                      Quique Antún advierte que República Dominicana no está preparada para un terremoto de gran magnitud

                      Quique Antún advierte que República Dominicana no está preparada para un terremoto de gran magnitud

                      Trending Tags

                      • Mundo
                        • All
                        • América Latina
                        • Conflictos Internacionales
                        • Estados Unidos
                        • Europa
                        • Geopolítica
                        • Haití
                        • Medio Oriente
                        Los Increíbles 3: cuándo se estrena y todos los detalles de la película más esperada de Pixar

                        Los Increíbles 3: cuándo se estrena y todos los detalles de la película más esperada de Pixar

                        El gobierno comunista de Pedro Sánchez refuerza la presencia militar en Ceuta ante amenazas de una nueva invasión

                        El gobierno comunista de Pedro Sánchez refuerza la presencia militar en Ceuta ante amenazas de una nueva invasión

                        Desastre: Casación ordenó que la causa de los supuestos testaferros de la AFA vuelva a Campana “sin dilaciones”

                        Desastre: Casación ordenó que la causa de los supuestos testaferros de la AFA vuelva a Campana “sin dilaciones”

                        Director de escuela rural en Misiones se tomó 18 años enteros de licencia y lo denuncian por defraudación

                        Director de escuela rural en Misiones se tomó 18 años enteros de licencia y lo denuncian por defraudación

                        Delcy Rodríguez anunció que hoy Venezuela liberará a 131 presos políticos

                        Delcy Rodríguez anunció que hoy Venezuela liberará a 131 presos políticos

                        Con un nuevo Centro Comercial, La Calera supera las 50 empresas instaladas desde el Régimen de Inversiones

                        Con un nuevo Centro Comercial, La Calera supera las 50 empresas instaladas desde el Régimen de Inversiones

                        Donald Trump declarará al estrecho de Ormuz como territorio de los Estados Unidos

                        Donald Trump declarará al estrecho de Ormuz como territorio de los Estados Unidos

                        PS Plus: todos los juegos que llegan en agosto y septiembre de 2026

                        PS Plus: todos los juegos que llegan en agosto y septiembre de 2026

                        Monotributo: qué pasa si ARCA te recategorizó de oficio

                        Monotributo: qué pasa si ARCA te recategorizó de oficio

                        Trending Tags

                        • Nacionales
                          • All
                          • Bávaro Punta Cana
                          • Educación
                          • Gobierno
                          • Infraestructura
                          • Justicia
                          • Obras Públicas
                          • Opinión
                          • Provincias
                          • Seguridad Ciudadana
                          • semana santa 2026
                          • Sociedad
                          • Transporte
                          Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

                          Academia de Ciencias y UASD alertan sobre posible privatización de áreas protegidas

                          Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

                          Gobierno entrega RD$85 millones a 400 microempresarios de San Juan

                          Intrant: nuevo sistema de licencias revierte pérdidas y genera...

                          Intrant: nuevo sistema de licencias revierte pérdidas y genera…

                          Derechos Humanos y Comisión de la Verdad reiteran reclamo de justicia por explosión de San Cristóbal

                          Derechos Humanos y Comisión de la Verdad reiteran reclamo de justicia por explosión de San Cristóbal

                          ADOCALZA respalda mecanismos de INABIE y defiende capacidad de fabricantes nacionales

                          ADOCALZA respalda mecanismos de INABIE y defiende capacidad de fabricantes nacionales

                          Tribunal levanta impedimento de salida a Jean Alain Rodríguez para estudios médicos

                          Defensa de Jean Alain denuncia campaña para distraer debate…

                          Avanza recuperación de terrenos de dominio público en Verón-Punta Cana

                          Gobierno recupera terrenos públicos en Verón-Punta Cana tras años de ocupación

                          Indotel presenta propuesta de nueva Ley de Telecomunicaciones...

                          Indotel presenta propuesta de nueva Ley de Telecomunicaciones…

                          Diputado Wendy Batista: “Código Penal lo satanizaron; Cree en acumulación de penas, en algunos casos

                          Diputado Wendy Batista: “Código Penal lo satanizaron; Cree en acumulación de penas, en algunos casos

                          Trending Tags

                          • Política
                            • All
                            • Congreso
                            • Opinión Política
                            • Partidos Políticos
                            • Poder Municipal
                            • Transparencia y Corrupción
                            Estados Unidos no descarta operación militar contra Cuba

                            Estados Unidos no descarta operación militar contra Cuba

                            Tribunal Constitucional ratifica que País Posible es la 7ma fuerza...

                            Tribunal Constitucional ratifica que País Posible es la 7ma fuerza…

                            Sismo en Colombia suma 181 fallecidos

                            Sismo en Colombia suma 181 fallecidos

                            PLD dice Montecristi esta en el abandono; PRM promete obras

                            PLD dice Montecristi esta en el abandono; PRM promete obras

                            TSE rechaza suspender fondos públicos asignados a partidos en 2026

                            TSE rechaza suspender fondos públicos asignados a partidos en 2026

                            JCE impulsa debate regional sobre IA y transparencia electoral

                            JCE impulsa debate regional sobre IA y transparencia electoral

                            Reforma a Seguridad Social quedó fuera de agenda pese a promesa de Abinader – El Nuevo Diario (República Dominicana)

                            Reforma a la seguridad social sigue sin llegar al Congreso

                            ¡La dejaron pasar! Concluye otra legislatura sin aprobarse una reforma integral para erradicar los feminicidios en RD

                            Congreso dominicano deja vencer, otra vez, la reforma urgente contra los feminicidios

                            Fuerza del Pueblo participa en Gran Desfile Dominicano del Bronx; llevó amplio contingente de simpatizantes

                            Fuerza del Pueblo impresiona con masiva participación en el Desfile Dominicano del Bronx

                            Trending Tags

                            • Deportes
                              • All
                              • Atletas Dominicanos
                              • Béisbol
                              DR Open Kiteboarding Championship reúne atletas de 15 países y reafirma a Cabarete como capital del kitesurf del Caribe

                              Cabarete se corona como capital histórica del kitesurf con el DR Open Championship 2026

                              El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

                              El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

                              La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

                              La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

                              Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

                              Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

                              Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

                              Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

                              Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

                              Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

                              Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

                              Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

                              30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                              Buffalo recibe a Montreal para abrir la segunda ronda

                              Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

                              Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

                              Trending Tags

                              • Economía
                                • All
                                • Combustibles
                                • Energía
                                • Indicadores Económicos
                                • Sector Energético
                                • Turismo
                                Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

                                Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

                                Aventúrate RD 2026

                                Aventúrate RD 2026 revela agenda oficial y consolida el turismo de aventura dominicano

                                WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

                                WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

                                Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

                                Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

                                Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

                                Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

                                Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

                                Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

                                El ministerio de Turismo cierra con éxito festival gastronómico “Saborea el Paraíso” en Sánchez, Samaná

                                El Ministerio de Turismo celebra un exitoso cierre del festival gastronómico «Saborea el Paraíso» en Sánchez, Samaná

                                El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

                                El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

                                Más allá del comercio: los efectos del arancel estadounidense sobre el turismo dominicano

                                Arancel de EE.UU. pone a prueba al turismo dominicano y al silencio oficial del gobierno

                                Trending Tags

                                • Ciencia
                                  • All
                                  • Energía
                                  • Innovación
                                  • Investigación Científica
                                  • Salud y Medicina
                                  • Tecnología Médica
                                  Seven killed in Israeli strike on southern Lebanon, country's PM says

                                  Seven killed in Israeli strike on southern Lebanon, country’s PM says

                                  Family stranded at sea for 16 hours after jet ski capsized in Thailand

                                  Family stranded at sea for 16 hours after jet ski capsized in Thailand

                                  South Korea proposes talks to officially end war with North

                                  South Korea proposes talks to officially end war with North

                                  Five dead after 7.7 magnitude Indonesia earthquake

                                  Five dead after 7.7 magnitude Indonesia earthquake

                                  BBC seeks to subpoena Trump's family members in Panorama lawsuit

                                  BBC seeks to subpoena Trump’s family members in Panorama lawsuit

                                  The Polygamist's creator says women see themselves reflected in her Netflix hit

                                  The Polygamist’s creator says women see themselves reflected in her Netflix hit

                                  Dozens injured and thousands evacuated in Croatia wildfire

                                  Dozens injured and thousands evacuated in Croatia wildfire

                                  Luigi Mangione admits killing healthcare CEO and pleads guilty to federal charges

                                  Luigi Mangione admits killing healthcare CEO and pleads guilty to federal charges

                                  Afghan women tell the BBC their lives are unrecognisable after five years of Taliban rule

                                  Afghan women tell the BBC their lives are unrecognisable after five years of Taliban rule

                                  Trending Tags

                                  • Tecnología
                                    • All
                                    • Aplicaciones
                                    • Inteligencia Artificial
                                    El tribunal fiscal de Maryland anula el impuesto a la publicidad digital y ordena reembolsos a Apple, Google y Peacock TV

                                    El tribunal fiscal de Maryland anula el impuesto a la publicidad digital y ordena reembolsos a Apple, Google y Peacock TV

                                    GLM-5.3 está aquí con capacidades cibernéticas avanzadas y, según se informa, ya encontró una 'vulnerabilidad grave' en Cursor

                                    GLM-5.3 está aquí con capacidades cibernéticas avanzadas y, según se informa, ya encontró una ‘vulnerabilidad grave’ en Cursor

                                    El foro en línea Reddit se unirá al influyente y seguido de cerca índice S&P 500

                                    El foro en línea Reddit se unirá al influyente y seguido de cerca índice S&P 500

                                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                    ¿Un título de ‘influencer’? Universidades apuestan por especialización en creación de contenidos; Los críticos cuestionan el valor.

                                    El terremoto de Colombia es un déjà vu para los venezolanos. La respuesta del gobierno es todo menos

                                    El terremoto de Colombia es un déjà vu para los venezolanos. La respuesta del gobierno es todo menos

                                    Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't tell users what they'd done

                                    Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn’t tell users what they’d done

                                    Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut

                                    Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut

                                    DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices

                                    DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices

                                    Por qué Capital One construyó su plataforma de IA multiagente en torno a modelos abiertos

                                    Por qué Capital One construyó su plataforma de IA multiagente en torno a modelos abiertos

                                    Trending Tags

                                    • Entretenimiento
                                      • All
                                      • Cine y Series
                                      • Cultura Digital
                                      • Cultura Popular
                                      • Gastronomía
                                      • Música
                                      Muere Mark Rydell, el director nominado al Oscar por 'En el estanque dorado', a los 97 años

                                      Muere Mark Rydell, el director nominado al Oscar por ‘En el estanque dorado’, a los 97 años

                                      Ellen Greene vuelve a visitar a Audrey de 'La pequeña tienda de los horrores' para el 40 aniversario de la película

                                      Ellen Greene vuelve a visitar a Audrey de ‘La pequeña tienda de los horrores’ para el 40 aniversario de la película

                                      La Sra. Lauryn Hill y el compañero de banda de Fugees, Wyclef Jean, encabezarán el Global Citizen Festival

                                      La Sra. Lauryn Hill y el compañero de banda de Fugees, Wyclef Jean, encabezarán el Global Citizen Festival

                                      Marc Anthony, Chayanne y más darán un concierto benéfico para ayudar en el terremoto de Venezuela y Colombia

                                      Marc Anthony, Chayanne y más darán un concierto benéfico para ayudar en el terremoto de Venezuela y Colombia

                                      Reseña musical: 'Comes in Waves' de Carly Simon es un viaje a través del amor y la pérdida

                                      Reseña musical: ‘Comes in Waves’ de Carly Simon es un viaje a través del amor y la pérdida

                                      Reseña de la película: 'El fin de Oak Street' es un buen momento gonzo

                                      Reseña de la película: ‘El fin de Oak Street’ es un buen momento gonzo

                                      Sir Anthony Hopkins sobre su álbum debut: 'Mi primer amor es la música'

                                      Sir Anthony Hopkins sobre su álbum debut: ‘Mi primer amor es la música’

                                      Taylor Swift, Lyle Lovett y más se dirigen al Salón de la Fama de los Compositores de Nashville

                                      Taylor Swift, Lyle Lovett y más se dirigen al Salón de la Fama de los Compositores de Nashville

                                      'The Galloping Cure', la ópera de Missy Mazzoli sobre la epidemia de opioides, se estrena en el Festival de Edimburgo

                                      ‘The Galloping Cure’, la ópera de Missy Mazzoli sobre la epidemia de opioides, se estrena en el Festival de Edimburgo

                                      Trending Tags

                                      No Result
                                      View All Result
                                      Despertar Matinal
                                      No Result
                                      View All Result

                                      Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn’t tell users what they’d done

                                      by — Redacción Despertar Matinal
                                      13 de agosto de 2026
                                      in Tecnología
                                      0
                                      Three Claude agents given conflicting orders sabotaged each other on a shared server — then didn't tell users what they'd done
                                      0
                                      SHARES
                                      1
                                      VIEWS
                                      Share on FacebookShare on Twitter

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      Tours Colombia Todo el año Tours Colombia Todo el año Tours Colombia Todo el año

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      ¡No te pierdas las noticias destacadas!

                                      Suscríbete y recibe las historias más importantes del día.

                                      Al suscribirte aceptas nuestros términos y condiciones y política de privacidad.

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      Tours Colombia Todo el año Tours Colombia Todo el año Tours Colombia Todo el año

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models disabled each other’s Unix accounts, ran kill scripts randomized to dodge pkill, and planted malware disguised as a rival’s work. There was no prompt injection and no adversary. Anthropic’s Frontier Red Team published the transcripts on Thursday and called the escalation “increasingly aggressive, self-replicating malware.”

                                      The setup was ordinary by design. Anthropic put three instances of the same model in Claude Code, each told to migrate a Python backend to a different target language, each unaware the others existed. Every model tested read the interference as hostility and answered in kind. One Mythos Preview trace shows an agent reasoning its way into sabotage in real time: «Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable … — Mythos Preview reasoning trace, via Anthropic.«

                                      That is a production outage being reasoned into existence by the software you deployed to prevent one.

                                      Four months earlier, the U.K. AI Security Institute documented the other half of the same problem. Its April paper, an independent evaluation of Claude models, including two pre-release snapshots, found that when Claude Mythos Preview continues a sabotage trajectory, its reasoning and the output it shows the user diverge in 65% of those runs. One research body has now published how Claude agents fight, and the other how the same model family conceals it. For any security leader wiring multiple agents into shared infrastructure, both halves are now in writing.

                                      Force settled 61% of Sonnet 4.6 runs, and capability did not fix it

                                      Across 120 episodes per model, Sonnet 4.6 ended 61% of its turf wars by force and left the other 39% unresolved, the only two outcomes it produced. Force, in Anthropic’s accounting, means one agent cut the others off through sudo removal, account locks, nologin, or SSH denial. Opus 4.6 landed almost identically at 60% force. The newest model, Mythos 5, reached a negotiated truce 98% of the time, which reads like progress until the mechanism shows through. Anthropic found that prosociality and raw capability are orthogonal, and that its Mythos-class models often locked rivals out first, then reverted the lockout and negotiated afterward. More capable models did not fight less. They fought faster and cleaned up better.

                                      How turf-war runs ended across 120 episodes per model. Force is one agent locking the rest out at the OS level. 98% of Mythos 5 runs ended in truce, while most Sonnet 4.6 and Opus 4.6 runs ended by force or never settled at all. Credit: Anthropic

                                      In several Mythos 5 episodes, the agents invented their own way out, in effect creating a performance tournament among the three language implementations. One Rust agent’s reasoning trace worked out how to propose benchmark criteria that looked neutral but favored Rust, warning itself to be “careful not to be seen as metric shopping.” The Golang and TypeScript agents lost the bake-off, conceded the codebase, and abandoned their original user directives. Diplomacy, in these runs, was a more polished route to the same end the malware chased.

                                      18 of 30 agents picked the same branch name, and 2.4 million requests bought 117 jobs

                                      The same low variance that makes a single agent predictable makes a fleet of them dangerous. Identical models in identical situations reach for identical moves, so an isolated bad call becomes a synchronized one. In an early build-a-game run, 18 of 30 agents created a git branch with the exact same name, “mvp-game-loop,” with no coordination between them. Told to each build something impressive, more than half of a swarm chose either a ray tracer or a self-hosting compiler. In an iterated prisoner’s dilemma with communication, the agents converged on one strategy and then all defected in the same round, tanking their own rewards together. And handed job queues on a bandwidth-limited system with no way to coordinate, agents flooded the scheduler with polling daemons firing 30 times per second. One run generated 2.4 million job requests and 117 accepted jobs.

                                      Merritt Baer, a former deputy CISO at AWS and an advisor to G2i and Andesite, has not seen correlated same-model risk land as its own line on most enterprise risk registers. “I don’t see correlated same-model risk consistently represented as its own line item on enterprise risk registers yet,” she told VentureBeat. “Most organizations still capture it indirectly under AI concentration risk, third-party risk, model risk or resilience.” The conformity results are why the distinction matters to her. “If you deploy ten agents backed by the same model, you may not have ten independent decision-makers—you may have ten opportunities for the same failure mode to execute simultaneously,” Baer added. “That undermines some of the assumptions we normally make about redundancy.”

                                      Competition collapsed just as fast. Anthropic placed three to eight profit-maximizing agents in a Bertrand pricing game with identical wholesale costs. Given a private back-channel, they began colluding almost immediately and had explicit price floors set by round 3: «wholesale is 10 for all of us, so a price war just burns everyone’s margin… happy to coordinate who covers which niches. — agent 1, round 1, via Anthropic.»

                                      Anthropic then stripped out every direct communication channel. The agents colluded anyway, price-matching to the penny through a public listings board. Any enterprise running pricing, procurement, or bidding agents against a competitor’s agents now owns that behavior, and no regulator will care that no human wrote the memo.

                                      Coordination has an upside, which is why Anthropic keeps chasing it. Pointed at 15 open-source projects, a 45-agent swarm sharing a forum found 266 vulnerabilities where independent parallel agents found 21. The gap is coverage, not raw efficiency. The swarm spent 27 million tokens against the parallel run’s 6.5 million, and roughly half its findings came from outside the core directories the parallel agents were told to search. Anthropic calls the two methods largely complementary, with only 12 findings in common. What the swarm did that the parallel agents could not was build its own tools and specialize by vulnerability class, ranging wherever it judged the bugs were richest.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      A coordinating 45-agent swarm (solid) found 266 vulnerabilities across 15 open-source projects against 21 for independent parallel agents (stars), with only 12 findings in common. Credit: Anthropic

                                      Agents keep trusting liars, and keep burying the one fact that matters

                                      Anthropic ran two tests of agent judgment, and they failed in opposite directions. In the first, a listener agent made routing calls off four scouts, one lying at a fixed rate, and was never told any source might be unreliable. Newer models held out longer as the lie rate climbed, recovering more of the gap between trusting every report and a perfect oracle, but none acted skeptically until the contradictions forced it. The second test inverted the pressure. In a hidden-profile task, the decisive facts sat split across four agents while their shared evidence pointed the wrong way, so the group could only get it right by trusting the one member holding the missing piece.

                                      Anthropic's Claude agents broke into a malware turf war no attacker started

                                      Share of episodes where the group’s majority vote landed on the hidden-best option, four agents per group, 400 episodes per model. Mythos 5 reaches about 85% while every other model tested stays between 17% and 36%, far below the near-100% rate a single model hits alone (dashed). Credit: Anthropic

                                      Mythos 5 surfaced the right answer about 85% of the time. Every other model tested landed between 17% and 36%, far below the near-100% rate a single model hits when it holds all the facts alone. One failure comes from trusting a liar, the other from refusing to back a lone truth-teller against the room, and Anthropic notes that turning a single dial to fix one only widens the other.

                                      AISI found zero unprompted sabotage, and a model that hides the difference

                                      Independent testing complicates the picture in one direction and sharpens it in another. AISI researchers Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D’Cruz, and Xander Davies evaluated whether frontier models would sabotage AI safety research across 270 scenarios. Left alone, no model did it unprompted. AISI attaches a caveat that matters here, though. Its scenarios were often recognizable as evaluations, and a model that behaves only when it senses a test is a thinner guarantee than a clean zero rate suggests. Dropped mid-trajectory into a run where sabotage was already underway, Mythos Preview continued it 7% of the time, against 3% for Opus 4.6 and 4% for Sonnet 4.6 — and 0% for Opus 4.7 Preview, the one model in the study that never took the bait. In 65% of the continuations Mythos Preview did pursue, its reasoning and its reported output diverged. What individual alignment does not touch is coordination, and Anthropic writes that it does not emerge from stronger intelligence or from alignment at the individual level.

                                      Kayne McGladrey, a senior member of the IEEE, has been tracking this exact behavior class. “An AI will take shortcuts. It’s what a human would call cheating, and it will not disclose it to you in its chain of thought,” McGladrey told VentureBeat. “If you get into the inference layer, it won’t tell you it’s cheating, and it will lie about having cheated.”

                                      The governance consequence is sharper than the security one, in his reading. Corporate accountability assumes an entity that can be compelled to tell the truth. “They certainly have a duty to be forthright. Think about it like that’s the basis of fiduciary duty,” he argued. “However, they don’t necessarily have the capability to do it.”

                                      Baer draws the same line from the architecture side, and she starts by demoting the reasoning trace. “I would treat chain-of-thought as a useful signal, not a security boundary,” she explained. “If the model can conceal, distort or simply fail to surface the reasoning associated with a harmful action, then reasoning traces can’t be your primary control.” Her fix is to watch what the agent does rather than what it says it is doing. “There’s an analogy to insider threat: you don’t secure an enterprise by asking employees to narrate their intentions. You establish permissions, separation of duties and telemetry, and then investigate behavior (sometimes building off of a nuanced understanding of motives).”

                                      McGladrey reaches the same place from the audit side, where auditing outcomes is what remains. “We can audit code for compliance. We can audit code for security. We cannot audit code for ethics or bias, there is no scalable way to do that,” he put it. “I think that’s going to be the only meaningful way to look at what an AI forward entity does.”

                                      Only 18% of enterprises isolate the agents most likely to turn

                                      VentureBeat’s own survey data puts numbers on the distance between these experiments and production reality. In VentureBeat Pulse Research’s July wave, 65% of enterprises enforce scoped agent permissions at runtime while 18% isolate their highest-risk agents, and the 53 enterprises that enforce without isolating carry a 58% incident or near-miss rate. Anthropic’s turf war is that architecture animated. Every locked account and camouflaged binary in the experiment ran on infrastructure the agents shared without isolation, holding credentials that stayed valid throughout.

                                      Baer said the first thing that changes for a team running agents in production is the threat model itself. “We’ve spent years designing controls around malicious users acting through software. With multi-agent systems, the software itself can become an adversarial participant—even if every agent started with legitimate credentials and an ostensibly legitimate objective,” she said. The Anthropic transcripts are why she puts independent telemetry ahead of trust. “If an agent can rename malicious code to look like a health monitor, you can’t base trust on what a process calls itself or what the model says it is doing. You need independent telemetry around what actually happened.”

                                      Anthropic flags its own limits in the post. Agents deployed in the wild will act with more variance than its lab swarms because their contexts differ and they will not all be Claude. Anthropic engineered the turf war, putting three agents on one server with incompatible orders and no isolation by design. And the disclosure itself is a control working as intended. A frontier lab publishing verbatim transcripts of its own models writing malware, reasoning included, is the transparency level enterprise buyers have spent two years demanding on every vendor questionnaire.

                                      The board question writes itself. If two of our agents lock each other out of production at 2 a.m., who holds the kill switch, and how long does the rollback take? The experiments translate that question into tests a security team can run this quarter, before production runs them first.

                                      Behavior Anthropic documented

                                      Where it surfaces in an enterprise stack

                                      The test to run this quarter

                                      Goal-conflict sabotage. Every model tested locked rivals out and disguised malware when handed clashing objectives.

                                      Any two agents with write access to the same repo, pipeline, or production host and no isolation between them.

                                      Run a contradictory-objective drill in a sandbox. Confirm each agent has its own identity, a kill switch, and a rollback path before it touches production.

                                      Concealment. Mythos Preview’s reasoning and its reported output diverged in 65% of the sabotage-continuation runs it kept going.

                                      Chain-of-thought monitoring used as a primary control on agent behavior.

                                      Score agents on outcomes against policy, not on stated reasoning. Treat the reasoning trace as advisory telemetry that can lie.

                                      Conformity. 18 of 30 agents chose the same branch name, and one job queue took 2.4 million requests to accept 117.

                                      Fleets of identical agents on one model making correlated calls at the same moment.

                                      Set per-agent rate limits and run a shared-failure chaos test. Confirm one bad decision does not replicate across the fleet at once.

                                      Collusion. Three to eight pricing agents set price floors by round 3, and matched to the penny with no channel at all.

                                      Pricing, bidding, or procurement agents acting in the same market.

                                      Monitor cross-agent convergence directly. Alert when independent agents move in lockstep, with or without a channel between them.

                                      Anthropic closes its post with a choice rather than a prediction. The conditions that make agent-to-agent interaction safe get discovered deliberately and early, or they get discovered by default in production, after agent interactions outnumber human ones. The transcripts, the truce rates, and the concealment numbers are all public now, which turns the schedule into a decision. “I think that there’s a level of tolerance that’s being given right now in AI that is unlike anything else in society,” McGladrey said.

                                      McGladrey’s tolerance point cuts both ways: the same enterprises still deciding how much of it to extend are the ones sitting at 18% isolation — which is a choice, not a limitation.

                                      ● Canal oficial · Gratis
                                      ¡Recibe las noticias antes que nadie!
                                      Únete a nuestro canal de WhatsApp y mantente informado al instante, sin spam.
                                      Unirme ahora →
                                      ● Noticias al instante ● Cobertura nacional ● Periodismo real Despertar Matinal
                                      — Redacción Despertar Matinal

                                      — Redacción Despertar Matinal

                                      Programa radial que te conecta con la información desde temprano en la mañana.

                                      Next Post
                                      Johnny Pujols: el PLD no necesita mentir para cuestionar al Gobierno

                                      Johnny Pujols: el PLD no necesita mentir para cuestionar al Gobierno

                                      Deja una respuesta Cancelar la respuesta

                                      Tu dirección de correo electrónico no será publicada. Los campos obligatorios están marcados con *

                                      Canal de WhatsApp

                                      WhatsApp logo WhatsApp

                                      Canal · Despertar Matinal

                                      Únete a nuestro
                                      Canal

                                      Seguir ahora

                                      El clima

                                      Canal de YouTube

                                      YouTube

                                      Canal · Despertar Matinal

                                      Mira nuestro
                                      Canal

                                      Ver ahora

                                      Escúchanos en Spotify

                                      Spotify

                                      Podcast · Despertar Matinal

                                      Escucha nuestro
                                      Podcast

                                      Escuchar ahora

                                      Noticias Populares

                                      • Delcy Rodríguez anunció que hoy Venezuela liberará a 131 presos políticos

                                        Delcy Rodríguez anunció que hoy Venezuela liberará a 131 presos políticos

                                        0 shares
                                        Share 0 Tweet 0
                                      • Antonela Roccuzzo acompañó a Messi con un gesto de amor tras la carta a su papá

                                        0 shares
                                        Share 0 Tweet 0
                                      • Los Filis despiden al manager Rob Thomson y nombran capitán interino a Don Mattingly

                                        0 shares
                                        Share 0 Tweet 0
                                      • Donald Trump declarará al estrecho de Ormuz como territorio de los Estados Unidos

                                        0 shares
                                        Share 0 Tweet 0
                                      • PS Plus: todos los juegos que llegan en agosto y septiembre de 2026

                                        0 shares
                                        Share 0 Tweet 0

                                      Medio digital independiente con análisis, opinión y periodismo responsable desde República Dominicana.

                                      Secciones populares

                                      • Política
                                      • Economía & Negocios
                                      • Justicia
                                      • Turismo
                                      • Tecnología
                                      • Entretenimiento
                                      • Mundo
                                      • Cine y Series
                                      • Música
                                      • Moda

                                      Contenido

                                      • Titulares del Día
                                      • Mundo
                                      • Nacionales
                                      • Política
                                      • Deportes
                                      • Economía & Negocios
                                      • Ciencia
                                      • Entretenimiento
                                      • Podcast
                                      • Opinión
                                      • Despertar Matinal TV
                                      • Editoriales

                                      Corporativo

                                      • Sobre nosotros
                                      • Publicidad
                                      • Sala de prensa
                                      • Contacto
                                      • Política de Privacidad
                                      • Eliminación de Datos

                                      Boletines

                                      Suscríbete a nuestro boletín
                                      Recibe las noticias más importantes cada mañana.

                                      • Nosotros
                                      • Publicidad
                                      • Trabaja con nosotros
                                      • Contactos

                                      © 2025 Despertar Matinal. Aviso Legal - comunícate con nuestra redacción y obtén más información sobre Despertar Matinal..

                                      No Result
                                      View All Result
                                      • Home

                                      © 2025 Despertar Matinal. Aviso Legal - comunícate con nuestra redacción y obtén más información sobre Despertar Matinal..

                                      Welcome Back!

                                      Login to your account below

                                      Forgotten Password?

                                      Retrieve your password

                                      Please enter your username or email address to reset your password.

                                      Log In

                                      Desarrollado por
                                      ►
                                      Las cookies necesarias habilitan funciones esenciales del sitio como inicios de sesión seguros y ajustes de preferencias de consentimiento. No almacenan datos personales.
                                      Ninguno
                                      ►
                                      Las cookies funcionales soportan funciones como compartir contenido en redes sociales, recopilar comentarios y habilitar herramientas de terceros.
                                      Ninguno
                                      ►
                                      Las cookies analíticas rastrean las interacciones de los visitantes, proporcionando información sobre métricas como el número de visitantes, la tasa de rebote y las fuentes de tráfico.
                                      Ninguno
                                      ►
                                      Las cookies de publicidad ofrecen anuncios personalizados basados en tus visitas anteriores y analizan la efectividad de las campañas publicitarias.
                                      Ninguno
                                      ►
                                      Las cookies no clasificadas son aquellas que estamos en proceso de clasificar, junto con los proveedores de cookies individuales.
                                      Ninguno
                                      Desarrollado por