• Nosotros
  • Publicidad
  • Trabaja con nosotros
  • Contactos
martes, agosto 18, 2026
  • Login
No Result
View All Result
NEWSLETTER
Despertar Matinal
  • Titulares del Día
    • All
    • En Portada
    Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

    Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

    Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

    Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

    Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

    Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

    Ito Bisonó asume como ministro de Relaciones Exteriores con una trayectoria de gestión pública y amplios vínculos internacionales

    Ito Bisonó asume como ministro de Relaciones Exteriores con una trayectoria de gestión pública y amplios vínculos internacionales

    Raúl Martínez: seis años bastan para exigir resultados

    Raúl Martínez: seis años bastan para exigir resultados

    Procurador fiscal pide aumento salarial para representantes del Ministerio Público

    Procurador fiscal pide aumento salarial para representantes del Ministerio Público

    El Instituto Duartiano aboga por preservar la autodeterminación de RD ante versiones sobre presiones de EE. UU.

    El Instituto Duartiano aboga por preservar la autodeterminación de RD ante versiones sobre presiones de EE. UU.

    Julito Fulcar asume la Vicepresidencia del Senado y coloca a Peravia en la dirección de la Cámara Alta

    Julito Fulcar asume la Vicepresidencia del Senado y coloca a Peravia en la dirección de la Cámara Alta

    PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

    PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

    Trending Tags

    • Mundo
      • All
      • América Latina
      • Conflictos Internacionales
      • Estados Unidos
      • Europa
      • Geopolítica
      • Haití
      • Medio Oriente
      Histórico: YPF habilitó la compra y venta de acciones desde su propia app

      Histórico: YPF habilitó la compra y venta de acciones desde su propia app

      La economía no se entiende mirando por el espejo retrovisor

      La economía no se entiende mirando por el espejo retrovisor

      Córdoba: el Nuevocentro Shopping gestiona obras de ampliación para incorporar firmas internacionales

      Córdoba: el Nuevocentro Shopping gestiona obras de ampliación para incorporar firmas internacionales

      Nueva estafa con ARCA: cómo funciona el correo falso que roba datos

      Nueva estafa con ARCA: cómo funciona el correo falso que roba datos

      El CEO de PlayStation habló sobre la fecha de lanzamiento de la PS6

      El CEO de PlayStation habló sobre la fecha de lanzamiento de la PS6

      Honor presentó el Robot Phone: cuánto cuesta y qué puede hacer

      Honor presentó el Robot Phone: cuánto cuesta y qué puede hacer

      Adidas promociona el islam en Argentina utilizando modelos con vestimenta islámica

      Adidas promociona el islam en Argentina utilizando modelos con vestimenta islámica

      El gobierno de Putin sentenció a 11 años de prisión a un político opositor por criticar la guerra en Ucrania

      El gobierno de Putin sentenció a 11 años de prisión a un político opositor por criticar la guerra en Ucrania

      La verdadera razón de la suba de la mora: La inflación ya no licúa las deudas y la tarjeta de crédito

      La verdadera razón de la suba de la mora: La inflación ya no licúa las deudas y la tarjeta de crédito

      Trending Tags

      • Nacionales
        • All
        • Bávaro Punta Cana
        • Educación
        • Gobierno
        • Infraestructura
        • Justicia
        • Obras Públicas
        • Opinión
        • Provincias
        • Seguridad Ciudadana
        • semana santa 2026
        • Sociedad
        • Transporte
        Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

        Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

        Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

        Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

        Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

        Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

        CRR Las Parras crea talleres industriales de producción de colchones, ropa y tapicería

        CRR Las Parras crea talleres industriales de producción de colchones, ropa y tapicería

        Alejandro Campos es juramentado por Eduardo Estrella como...

        Alejandro Campos es juramentado por Eduardo Estrella como…

        Ito Bisonó asume como ministro de Relaciones Exteriores con una trayectoria de gestión pública y amplios vínculos internacionales

        Ito Bisonó asume como ministro de Relaciones Exteriores con una trayectoria de gestión pública y amplios vínculos internacionales

        Raúl Martínez: seis años bastan para exigir resultados

        Raúl Martínez: seis años bastan para exigir resultados

        Procurador fiscal pide aumento salarial para representantes del Ministerio Público

        Procurador fiscal pide aumento salarial para representantes del Ministerio Público

        Intrant logra por primera vez la triple certificación ISO en gestión...

        Milton Morrison concluye histórica gestión en el Intrant

        Trending Tags

        • Política
          • All
          • Congreso
          • Opinión Política
          • Partidos Políticos
          • Poder Municipal
          • Transparencia y Corrupción
          Milton Morrison reafirma alianza con Abinader y anuncia nueva...

          Milton Morrison reafirma alianza con Abinader y anuncia nueva…

          Empresarios de Hato Mayor expresan respaldo a Leonel Fernández y fortalecen proyecto político rumbo a 2028

          Empresarios de Hato Mayor expresan respaldo a Leonel Fernández y fortalecen proyecto político rumbo a 2028

          PRM en Santo Domingo Norte resalta gestión del presidente...

          PRM en Santo Domingo Norte resalta gestión del presidente…

          ARTICULO: De los millones de seguidores al poder: gobernar un país no es hacer un reality en YouTube

          ARTICULO: De los millones de seguidores al poder: gobernar un país no es hacer un reality en YouTube

          Estados Unidos no descarta operación militar contra Cuba

          Estados Unidos no descarta operación militar contra Cuba

          Tribunal Constitucional ratifica que País Posible es la 7ma fuerza...

          Tribunal Constitucional ratifica que País Posible es la 7ma fuerza…

          Sismo en Colombia suma 181 fallecidos

          Sismo en Colombia suma 181 fallecidos

          PLD dice Montecristi esta en el abandono; PRM promete obras

          PLD dice Montecristi esta en el abandono; PRM promete obras

          TSE rechaza suspender fondos públicos asignados a partidos en 2026

          TSE rechaza suspender fondos públicos asignados a partidos en 2026

          Trending Tags

          • Deportes
            • All
            • Atletas Dominicanos
            • Béisbol
            DR Open Kiteboarding Championship reúne atletas de 15 países y reafirma a Cabarete como capital del kitesurf del Caribe

            Cabarete se corona como capital histórica del kitesurf con el DR Open Championship 2026

            El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

            El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

            La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

            La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

            Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

            Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

            Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

            Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

            Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

            Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

            Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

            Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

            30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

            Buffalo recibe a Montreal para abrir la segunda ronda

            Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

            Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

            Trending Tags

            • Economía
              • All
              • Combustibles
              • Energía
              • Indicadores Económicos
              • Sector Energético
              • Turismo
              Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

              Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

              Aventúrate RD 2026

              Aventúrate RD 2026 revela agenda oficial y consolida el turismo de aventura dominicano

              WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

              WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

              Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

              Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

              Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

              Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

              Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

              Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

              El ministerio de Turismo cierra con éxito festival gastronómico “Saborea el Paraíso” en Sánchez, Samaná

              El Ministerio de Turismo celebra un exitoso cierre del festival gastronómico «Saborea el Paraíso» en Sánchez, Samaná

              El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

              El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

              Más allá del comercio: los efectos del arancel estadounidense sobre el turismo dominicano

              Arancel de EE.UU. pone a prueba al turismo dominicano y al silencio oficial del gobierno

              Trending Tags

              • Ciencia
                • All
                • Energía
                • Innovación
                • Investigación Científica
                • Salud y Medicina
                • Tecnología Médica
                Nigeria influencer 'KC Luxury' arrested after cocaine destined for the UK seized

                Nigeria influencer ‘KC Luxury’ arrested after cocaine destined for the UK seized

                Pakistan court orders ex-PM Imran Khan be moved to hospital from jail

                Pakistan court orders ex-PM Imran Khan be moved to hospital from jail

                "El sobreviviente": Collado es el único ministro que se mantiene

                «El sobreviviente»: Collado es el único ministro que se mantiene

                Lake Kariba disaster: Death toll from Zimbabwe ferry accident rises to 93

                Lake Kariba disaster: Death toll from Zimbabwe ferry accident rises to 93

                Zambia elections: Hakainde Hichilema re-elected president as main rival goes into hiding over alleged threats

                Zambia elections: Hakainde Hichilema re-elected president as main rival goes into hiding over alleged threats

                Hayden Panettiere: Ex Wladimir Klitschko says family is in 'profound shock and grief'

                Hayden Panettiere: Ex Wladimir Klitschko says family is in ‘profound shock and grief’

                IVF staff accused of misleading UK parents about donors at northern Cyprus clinics

                IVF staff accused of misleading UK parents about donors at northern Cyprus clinics

                Jharkhand protests: Devendra Nath Mahto ends hunger strike after 16 days

                Jharkhand protests: Devendra Nath Mahto ends hunger strike after 16 days

                Zhu Rongji: China censors public mourning as it holds former premier's funeral

                Zhu Rongji: China censors public mourning as it holds former premier’s funeral

                Trending Tags

                • Tecnología
                  • All
                  • Aplicaciones
                  • Inteligencia Artificial
                  85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one

                  85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one

                  El Reino Unido y Google prueban cambios en las rutas de vuelo para abordar el impacto climático de la aviación

                  El Reino Unido y Google prueban cambios en las rutas de vuelo para abordar el impacto climático de la aviación

                  La IA del comercio se está fragmentando. He aquí por qué eso es importante.

                  La IA del comercio se está fragmentando. He aquí por qué eso es importante.

                  Las empresas están pagando de más por consultas simples de IA: la puerta de enlace de Snowflake ahora se enruta automáticamente para reducir los costos hasta 3 veces

                  Las empresas están pagando de más por consultas simples de IA: la puerta de enlace de Snowflake ahora se enruta automáticamente para reducir los costos hasta 3 veces

                  OpenAI lanza ChatGPT para adolescentes, prometiendo un chatbot más apropiado para la edad

                  OpenAI lanza ChatGPT para adolescentes, prometiendo un chatbot más apropiado para la edad

                  Las tasas de vacunación escolar en EE. UU. vuelven a caer y las exenciones alcanzan un nivel récord

                  Las tasas de vacunación escolar en EE. UU. vuelven a caer y las exenciones alcanzan un nivel récord

                  Qwen3.8-27B ejecuta agentes de codificación de vanguardia y razonamiento local, no se requiere API en la nube

                  Qwen3.8-27B ejecuta agentes de codificación de vanguardia y razonamiento local, no se requiere API en la nube

                  30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                  Discord detiene las transmisiones en vivo en Brasil después de que un organismo de control citara fallas en la seguridad infantil

                  Cursor launches Origin code hosting platform as GitHub outage exposes opening in AI coding race

                  Cursor launches Origin code hosting platform as GitHub outage exposes opening in AI coding race

                  Trending Tags

                  • Entretenimiento
                    • All
                    • Cine y Series
                    • Cultura Digital
                    • Cultura Popular
                    • Gastronomía
                    • Música
                    El cofundador de ESPN, Bill Rasmussen, muere a los 93 años por los efectos de la enfermedad de Parkinson

                    El cofundador de ESPN, Bill Rasmussen, muere a los 93 años por los efectos de la enfermedad de Parkinson

                    Fox Sports transmitirá 35 partidos de voleibol femenino, incluidos 8 en Fox

                    Fox Sports transmitirá 35 partidos de voleibol femenino, incluidos 8 en Fox

                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                    Una película de animación china calificada de «terrible» se convierte en un éxito de taquilla

                    Shakira realiza visita sorpresa a Colombia afectada por el terremoto y se compromete a construir nuevas escuelas

                    Shakira realiza visita sorpresa a Colombia afectada por el terremoto y se compromete a construir nuevas escuelas

                    Bonnie Tyler es recordada como estrella mundial en su funeral en Gales

                    Bonnie Tyler es recordada como estrella mundial en su funeral en Gales

                    Dave Marsh, biógrafo y crítico musical de Bruce Springsteen, muere a los 76 años

                    Dave Marsh, biógrafo y crítico musical de Bruce Springsteen, muere a los 76 años

                    Andrew Garfield encuentra maravillas en la vida cotidiana en 'El árbol mágico lejano'

                    Andrew Garfield encuentra maravillas en la vida cotidiana en ‘El árbol mágico lejano’

                    Taquilla: 'Spider-Man' se mantiene en la cima mientras dos películas de dinosaurios luchan por el tercer puesto

                    Taquilla: ‘Spider-Man’ se mantiene en la cima mientras dos películas de dinosaurios luchan por el tercer puesto

                    Muere Bou Meng, artista camboyano y superviviente de un centro de tortura de los Jemeres Rojos, a los 85 años

                    Muere Bou Meng, artista camboyano y superviviente de un centro de tortura de los Jemeres Rojos, a los 85 años

                    Trending Tags

                    • Titulares del Día
                      • All
                      • En Portada
                      Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

                      Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

                      Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

                      Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

                      Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

                      Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

                      Ito Bisonó asume como ministro de Relaciones Exteriores con una trayectoria de gestión pública y amplios vínculos internacionales

                      Ito Bisonó asume como ministro de Relaciones Exteriores con una trayectoria de gestión pública y amplios vínculos internacionales

                      Raúl Martínez: seis años bastan para exigir resultados

                      Raúl Martínez: seis años bastan para exigir resultados

                      Procurador fiscal pide aumento salarial para representantes del Ministerio Público

                      Procurador fiscal pide aumento salarial para representantes del Ministerio Público

                      El Instituto Duartiano aboga por preservar la autodeterminación de RD ante versiones sobre presiones de EE. UU.

                      El Instituto Duartiano aboga por preservar la autodeterminación de RD ante versiones sobre presiones de EE. UU.

                      Julito Fulcar asume la Vicepresidencia del Senado y coloca a Peravia en la dirección de la Cámara Alta

                      Julito Fulcar asume la Vicepresidencia del Senado y coloca a Peravia en la dirección de la Cámara Alta

                      PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

                      PRD califica de desastrosa gestión de Abinader y afirma que el país ha retrocedido

                      Trending Tags

                      • Mundo
                        • All
                        • América Latina
                        • Conflictos Internacionales
                        • Estados Unidos
                        • Europa
                        • Geopolítica
                        • Haití
                        • Medio Oriente
                        Histórico: YPF habilitó la compra y venta de acciones desde su propia app

                        Histórico: YPF habilitó la compra y venta de acciones desde su propia app

                        La economía no se entiende mirando por el espejo retrovisor

                        La economía no se entiende mirando por el espejo retrovisor

                        Córdoba: el Nuevocentro Shopping gestiona obras de ampliación para incorporar firmas internacionales

                        Córdoba: el Nuevocentro Shopping gestiona obras de ampliación para incorporar firmas internacionales

                        Nueva estafa con ARCA: cómo funciona el correo falso que roba datos

                        Nueva estafa con ARCA: cómo funciona el correo falso que roba datos

                        El CEO de PlayStation habló sobre la fecha de lanzamiento de la PS6

                        El CEO de PlayStation habló sobre la fecha de lanzamiento de la PS6

                        Honor presentó el Robot Phone: cuánto cuesta y qué puede hacer

                        Honor presentó el Robot Phone: cuánto cuesta y qué puede hacer

                        Adidas promociona el islam en Argentina utilizando modelos con vestimenta islámica

                        Adidas promociona el islam en Argentina utilizando modelos con vestimenta islámica

                        El gobierno de Putin sentenció a 11 años de prisión a un político opositor por criticar la guerra en Ucrania

                        El gobierno de Putin sentenció a 11 años de prisión a un político opositor por criticar la guerra en Ucrania

                        La verdadera razón de la suba de la mora: La inflación ya no licúa las deudas y la tarjeta de crédito

                        La verdadera razón de la suba de la mora: La inflación ya no licúa las deudas y la tarjeta de crédito

                        Trending Tags

                        • Nacionales
                          • All
                          • Bávaro Punta Cana
                          • Educación
                          • Gobierno
                          • Infraestructura
                          • Justicia
                          • Obras Públicas
                          • Opinión
                          • Provincias
                          • Seguridad Ciudadana
                          • semana santa 2026
                          • Sociedad
                          • Transporte
                          Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

                          Diputados de la FP someten resolución para interpelar al ministro de Educación ante deterioro del sistema educativo a pocos días del inicio del año escolar 2026-2027

                          Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

                          Harold Modesto: “Ministerio Público influenció en cambios nuevo CP”; Pide abogados a estudiarlo y clama a aplicarlo bien”

                          Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

                          Méndez asume dirección del INTRANT con firme convicción de hacer cumplir la ley

                          CRR Las Parras crea talleres industriales de producción de colchones, ropa y tapicería

                          CRR Las Parras crea talleres industriales de producción de colchones, ropa y tapicería

                          Alejandro Campos es juramentado por Eduardo Estrella como...

                          Alejandro Campos es juramentado por Eduardo Estrella como…

                          Ito Bisonó asume como ministro de Relaciones Exteriores con una trayectoria de gestión pública y amplios vínculos internacionales

                          Ito Bisonó asume como ministro de Relaciones Exteriores con una trayectoria de gestión pública y amplios vínculos internacionales

                          Raúl Martínez: seis años bastan para exigir resultados

                          Raúl Martínez: seis años bastan para exigir resultados

                          Procurador fiscal pide aumento salarial para representantes del Ministerio Público

                          Procurador fiscal pide aumento salarial para representantes del Ministerio Público

                          Intrant logra por primera vez la triple certificación ISO en gestión...

                          Milton Morrison concluye histórica gestión en el Intrant

                          Trending Tags

                          • Política
                            • All
                            • Congreso
                            • Opinión Política
                            • Partidos Políticos
                            • Poder Municipal
                            • Transparencia y Corrupción
                            Milton Morrison reafirma alianza con Abinader y anuncia nueva...

                            Milton Morrison reafirma alianza con Abinader y anuncia nueva…

                            Empresarios de Hato Mayor expresan respaldo a Leonel Fernández y fortalecen proyecto político rumbo a 2028

                            Empresarios de Hato Mayor expresan respaldo a Leonel Fernández y fortalecen proyecto político rumbo a 2028

                            PRM en Santo Domingo Norte resalta gestión del presidente...

                            PRM en Santo Domingo Norte resalta gestión del presidente…

                            ARTICULO: De los millones de seguidores al poder: gobernar un país no es hacer un reality en YouTube

                            ARTICULO: De los millones de seguidores al poder: gobernar un país no es hacer un reality en YouTube

                            Estados Unidos no descarta operación militar contra Cuba

                            Estados Unidos no descarta operación militar contra Cuba

                            Tribunal Constitucional ratifica que País Posible es la 7ma fuerza...

                            Tribunal Constitucional ratifica que País Posible es la 7ma fuerza…

                            Sismo en Colombia suma 181 fallecidos

                            Sismo en Colombia suma 181 fallecidos

                            PLD dice Montecristi esta en el abandono; PRM promete obras

                            PLD dice Montecristi esta en el abandono; PRM promete obras

                            TSE rechaza suspender fondos públicos asignados a partidos en 2026

                            TSE rechaza suspender fondos públicos asignados a partidos en 2026

                            Trending Tags

                            • Deportes
                              • All
                              • Atletas Dominicanos
                              • Béisbol
                              DR Open Kiteboarding Championship reúne atletas de 15 países y reafirma a Cabarete como capital del kitesurf del Caribe

                              Cabarete se corona como capital histórica del kitesurf con el DR Open Championship 2026

                              El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

                              El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

                              La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

                              La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

                              Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

                              Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

                              Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

                              Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

                              Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

                              Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

                              Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

                              Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

                              30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                              Buffalo recibe a Montreal para abrir la segunda ronda

                              Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

                              Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

                              Trending Tags

                              • Economía
                                • All
                                • Combustibles
                                • Energía
                                • Indicadores Económicos
                                • Sector Energético
                                • Turismo
                                Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

                                Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

                                Aventúrate RD 2026

                                Aventúrate RD 2026 revela agenda oficial y consolida el turismo de aventura dominicano

                                WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

                                WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

                                Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

                                Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

                                Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

                                Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

                                Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

                                Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

                                El ministerio de Turismo cierra con éxito festival gastronómico “Saborea el Paraíso” en Sánchez, Samaná

                                El Ministerio de Turismo celebra un exitoso cierre del festival gastronómico «Saborea el Paraíso» en Sánchez, Samaná

                                El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

                                El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

                                Más allá del comercio: los efectos del arancel estadounidense sobre el turismo dominicano

                                Arancel de EE.UU. pone a prueba al turismo dominicano y al silencio oficial del gobierno

                                Trending Tags

                                • Ciencia
                                  • All
                                  • Energía
                                  • Innovación
                                  • Investigación Científica
                                  • Salud y Medicina
                                  • Tecnología Médica
                                  Nigeria influencer 'KC Luxury' arrested after cocaine destined for the UK seized

                                  Nigeria influencer ‘KC Luxury’ arrested after cocaine destined for the UK seized

                                  Pakistan court orders ex-PM Imran Khan be moved to hospital from jail

                                  Pakistan court orders ex-PM Imran Khan be moved to hospital from jail

                                  "El sobreviviente": Collado es el único ministro que se mantiene

                                  «El sobreviviente»: Collado es el único ministro que se mantiene

                                  Lake Kariba disaster: Death toll from Zimbabwe ferry accident rises to 93

                                  Lake Kariba disaster: Death toll from Zimbabwe ferry accident rises to 93

                                  Zambia elections: Hakainde Hichilema re-elected president as main rival goes into hiding over alleged threats

                                  Zambia elections: Hakainde Hichilema re-elected president as main rival goes into hiding over alleged threats

                                  Hayden Panettiere: Ex Wladimir Klitschko says family is in 'profound shock and grief'

                                  Hayden Panettiere: Ex Wladimir Klitschko says family is in ‘profound shock and grief’

                                  IVF staff accused of misleading UK parents about donors at northern Cyprus clinics

                                  IVF staff accused of misleading UK parents about donors at northern Cyprus clinics

                                  Jharkhand protests: Devendra Nath Mahto ends hunger strike after 16 days

                                  Jharkhand protests: Devendra Nath Mahto ends hunger strike after 16 days

                                  Zhu Rongji: China censors public mourning as it holds former premier's funeral

                                  Zhu Rongji: China censors public mourning as it holds former premier’s funeral

                                  Trending Tags

                                  • Tecnología
                                    • All
                                    • Aplicaciones
                                    • Inteligencia Artificial
                                    85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one

                                    85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one

                                    El Reino Unido y Google prueban cambios en las rutas de vuelo para abordar el impacto climático de la aviación

                                    El Reino Unido y Google prueban cambios en las rutas de vuelo para abordar el impacto climático de la aviación

                                    La IA del comercio se está fragmentando. He aquí por qué eso es importante.

                                    La IA del comercio se está fragmentando. He aquí por qué eso es importante.

                                    Las empresas están pagando de más por consultas simples de IA: la puerta de enlace de Snowflake ahora se enruta automáticamente para reducir los costos hasta 3 veces

                                    Las empresas están pagando de más por consultas simples de IA: la puerta de enlace de Snowflake ahora se enruta automáticamente para reducir los costos hasta 3 veces

                                    OpenAI lanza ChatGPT para adolescentes, prometiendo un chatbot más apropiado para la edad

                                    OpenAI lanza ChatGPT para adolescentes, prometiendo un chatbot más apropiado para la edad

                                    Las tasas de vacunación escolar en EE. UU. vuelven a caer y las exenciones alcanzan un nivel récord

                                    Las tasas de vacunación escolar en EE. UU. vuelven a caer y las exenciones alcanzan un nivel récord

                                    Qwen3.8-27B ejecuta agentes de codificación de vanguardia y razonamiento local, no se requiere API en la nube

                                    Qwen3.8-27B ejecuta agentes de codificación de vanguardia y razonamiento local, no se requiere API en la nube

                                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                    Discord detiene las transmisiones en vivo en Brasil después de que un organismo de control citara fallas en la seguridad infantil

                                    Cursor launches Origin code hosting platform as GitHub outage exposes opening in AI coding race

                                    Cursor launches Origin code hosting platform as GitHub outage exposes opening in AI coding race

                                    Trending Tags

                                    • Entretenimiento
                                      • All
                                      • Cine y Series
                                      • Cultura Digital
                                      • Cultura Popular
                                      • Gastronomía
                                      • Música
                                      El cofundador de ESPN, Bill Rasmussen, muere a los 93 años por los efectos de la enfermedad de Parkinson

                                      El cofundador de ESPN, Bill Rasmussen, muere a los 93 años por los efectos de la enfermedad de Parkinson

                                      Fox Sports transmitirá 35 partidos de voleibol femenino, incluidos 8 en Fox

                                      Fox Sports transmitirá 35 partidos de voleibol femenino, incluidos 8 en Fox

                                      30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                      Una película de animación china calificada de «terrible» se convierte en un éxito de taquilla

                                      Shakira realiza visita sorpresa a Colombia afectada por el terremoto y se compromete a construir nuevas escuelas

                                      Shakira realiza visita sorpresa a Colombia afectada por el terremoto y se compromete a construir nuevas escuelas

                                      Bonnie Tyler es recordada como estrella mundial en su funeral en Gales

                                      Bonnie Tyler es recordada como estrella mundial en su funeral en Gales

                                      Dave Marsh, biógrafo y crítico musical de Bruce Springsteen, muere a los 76 años

                                      Dave Marsh, biógrafo y crítico musical de Bruce Springsteen, muere a los 76 años

                                      Andrew Garfield encuentra maravillas en la vida cotidiana en 'El árbol mágico lejano'

                                      Andrew Garfield encuentra maravillas en la vida cotidiana en ‘El árbol mágico lejano’

                                      Taquilla: 'Spider-Man' se mantiene en la cima mientras dos películas de dinosaurios luchan por el tercer puesto

                                      Taquilla: ‘Spider-Man’ se mantiene en la cima mientras dos películas de dinosaurios luchan por el tercer puesto

                                      Muere Bou Meng, artista camboyano y superviviente de un centro de tortura de los Jemeres Rojos, a los 85 años

                                      Muere Bou Meng, artista camboyano y superviviente de un centro de tortura de los Jemeres Rojos, a los 85 años

                                      Trending Tags

                                      No Result
                                      View All Result
                                      Despertar Matinal
                                      No Result
                                      View All Result

                                      85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one

                                      by — Redacción Despertar Matinal
                                      18 de agosto de 2026
                                      in Tecnología
                                      0
                                      85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one
                                      0
                                      SHARES
                                      0
                                      VIEWS
                                      Share on FacebookShare on Twitter

                                      Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows.

                                      In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior. Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month.

                                      Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers, essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once.

                                      The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much autonomy companies give agents and how well they can verify them. It is increasingly a gap between confidence in the evaluation layer and evidence that the layer is getting better at preventing failures.

                                      The most revealing split appears inside the July data.

                                      Of the enterprises that experienced an AI feature clear testing only to go on to disappoint a customer, 4% placed complete faith in automated checks. Of those that had detected no comparable incident, 24% expressed full confidence — a sixfold difference.

                                      It makes sense: those who experienced test-passing agents failing in live production are, unsurprisingly, more likely to doubt the automated checking process.

                                      Perhaps it makes sense then, that companies geared toward tackling this problem — like automated agent error monitoring and mitigation platform Raindrop.ai — are seeing the market transform wildly from just a few months ago.

                                      «We are seeing the great-decline of evals as we know them,» Raindrop CTO Ben Hylak told VentureBeat in a direct message. «The Fortune 100 are increasingly reducing eval sets and deprioritizing maintenance. As systems grow more complex (MCPs, subagents, etc.) it becomes impossible to fully enumerate the failure cases. Instead, they’re leaning on anomaly and issue detection solutions, both before and after production.»

                                      A directional finding, not a market census

                                      VentureBeat fielded the July wave among 108 people representing companies with workforces of at least 100. This is down from 157 respondents in June.

                                      Of the 108, 69% described themselves as final AI-buying authorities or people who recommend and influence those purchases. The sample skewed toward midsize organizations: 63% worked at companies with 100 to 2,499 employees.

                                      The findings should be read directionally. The survey is self-selected rather than a probability sample, and the burned-vs.-unburned splits cited throughout this piece rest on groups of 41 to 53 respondents, and other cross-tabs in the report range from 40 to 68.

                                      The industry mix also changed: technology and software participation declined nine points, ending at 14%, while the retail and consumer share added four points and ended at 19%.

                                      The report nevertheless identifies four month-to-month changes worth noticing: more respondents professing complete confidence, fewer naming poor real-world alignment, more choosing integration ease as the decisive buying factor, and Braintrust gaining primary-platform share.

                                      Confidence in automated evals improved, but outcomes stayed flat

                                      VentureBeat’s June research identified an enterprise evaluation gap: companies were granting agents more authority faster than they were developing reliable ways to test them.

                                      July preserves the key number from that first wave. Across 265 enterprise responses over the two months, the proportion reporting at least one test-approved system that disappointed customers stayed within a single percentage point: 50% in June and 49% in July.

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      This figure does not mean that 49% of all agent runs fail, or that any particular evaluation product has a 49% failure rate. The survey asks whether an organization experienced at least one customer-facing incident in the previous year after an AI feature passed its internal tests. Companies that deploy far more agents have more opportunities to encounter such an incident.

                                      But that limitation does not make the result less important. An internal evaluation serves as a release gate. If roughly half of surveyed organizations have seen that gate approve a system that later fails in front of customers, a passing score cannot be treated as proof of production reliability.

                                      The cross-tab reinforces the point. Ten of the 41 enterprises with no identified testing miss placed complete faith in automation.

                                      Only two of the 53 previously burned enterprises said the same. Confidence is strongest among respondents with the least evidence that the release gate can fail.

                                      The enterprises that got burned are moving faster toward zero-human deployment

                                      The counterintuitive finding is what companies do after an evaluation miss.

                                      Overall, 67% either let an agent push code or change a system without a person’s approval in certain low-risk cases, or are modifying their pipelines to support that practice during the coming year. That is unchanged from June. In July, 37% already permitted it in limited cases and another 30% were building toward it.

                                      Among enterprises where a test-approved system had disappointed a customer, however, 85% were pursuing that no-approval model, compared with 61% in the group reporting no comparable incident. Only 11% of burned respondents rejected end-to-end deployment automation for the years ahead, versus 24% of unburned respondents.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 chart

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026.

                                      It would be easy to read that as recklessness, but the data supports another plausible explanation: deployment maturity. Organizations running more agents, at higher volume and across more consequential workflows, are more likely both to encounter failures and to have the engineering infrastructure needed for automated deployment.

                                      The survey cannot establish which explanation dominates. It does establish that a customer-visible incident does not appear to stop the move toward autonomy. Respondents with firsthand proof that testing can miss defects are also moving most aggressively to let those tests authorize production changes.

                                      If their per-deployment failure rate remains constant while deployment volume rises, the total incident count could grow even without the percentage of affected companies increasing. The July data does not measure incident volume, so that remains a risk implied by the pattern rather than a measured outcome.

                                      The release gate is automated, but production quality monitoring still lags

                                      Pre-deployment evaluation and production monitoring answer different questions. An evaluation asks whether an agent appears ready to ship. Production monitoring asks what the agent is doing after release and whether its live outputs remain correct.

                                      Most companies in the July sample still emphasize whether the system functions, not whether the answer is correct.

                                      Among the 106 valid responses to this question, 26% used inline quality assertions — automated judges or guardrails checking live traffic for output-quality problems. Another 26% focused on transaction traces such as infrastructure spans, token usage and raw inputs and outputs, while 24% mainly tracked gateway metrics such as latency, errors and cost.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026 monitoring uptime vs. wrong answers

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026

                                      Trace and gateway data can reveal outages, slowdowns and broken requests. They may not flag a fluent, fast and confidently wrong answer. Grouped by the report according to what each architecture actually watches, half of respondents monitored whether an agent was functioning, while just over a quarter automatically monitored whether its production output was correct.

                                      The gap is sharpest among the 40 respondents already permitting no-approval deployment in limited cases. Only 28% of that group automatically checked the meaning and correctness of live answers. In other words, most enterprises that have eliminated a person from at least some release decisions have not installed automated semantic-quality monitoring as the production backstop.

                                      This is the clearest operational lesson in the data. A pre-deployment test suite and infrastructure observability are necessary, but they do not cover the same failure mode. Enterprises need a way to detect bad outputs after the agent begins interacting with real users, data and tools — especially when nobody reviews the deployment decision first.

                                      An independent agent-evaluation market begins to take shape

                                      The vendor data offers a more encouraging sign: enterprises are adding dedicated evaluation tools, and specialist platforms are gaining ground.

                                      OpenAI’s native evals and traces narrowly led as the primary platform at 18%, followed by Confident AI’s DeepEval at 17% and Braintrust at 15%.

                                      Anthropic’s Claude Console and Workbench held 12%, tied with organizations reporting no dedicated evaluation platform. Three options each held 6%: internally built tools, Promptfoo and LangSmith.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 market landscape

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Braintrust’s primary share increased from 8% in June to 15% in July, the biggest gain by one vendor and the one the report flags as statistically significant. DeepEval rose from 12% to 17%. Use of no purpose-built platform declined five points to 12%, although that smaller change does not by itself confirm a trend.

                                      Because many companies use more than one tool, the broader footprints are larger. OpenAI native evaluation appeared somewhere in 31% of stacks, DeepEval in 27%, Braintrust in 22% and Anthropic’s native tooling in 20%. Custom internal tooling reached 14%, while Weights & Biases Weave and open-source Langfuse each reached 11%.

                                      These are adoption figures, not product-performance scores. The survey does not establish that one vendor produces more reliable agents than another. Still, the results point toward evaluation becoming a distinct enterprise software layer rather than a loose collection of internal scripts or a feature used only inside a model provider’s platform.

                                      Purchasing priorities are changing with that market. The proportion choosing integration ease as the decisive factor climbed 12 points to 39%, displacing cost, which fell from 28% to 23%. Evaluation accuracy ranked second at 28%. The combined average for satisfaction, implementation simplicity and economic value was 3.9 out of five.

                                      The move from price toward integration suggests enterprises increasingly want a tool they can install into existing development and monitoring pipelines now. Yet their leading success metric remains evaluation consistency at 38%, followed by fewer failures and regressions at 20%. Buyers are selecting for fit while still judging results on repeatability.

                                      Switching intent also cooled: 56% still expected to add or replace a platform during the coming year, down from 64% in June.

                                      The proportion staying put increased eight points, reaching 44%. Together with specialist adoption, they suggest some buyers are moving from evaluation to implementation.

                                      Human review is becoming the hedge against automated misses

                                      The budget data reveals how enterprises are managing the contradiction between greater autonomy and unreliable evaluation.

                                      People-centered review workflows edged narrowly ahead of production observability as the most frequently cited area for increased investment, 31% to 30%.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 human review

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Automated evaluation pipelines ranked third at 19%, followed by testing for safety and policy compliance at 16%. Only 6% said their reliability and evaluation budget was not increasing.

                                      Among enterprises that had experienced a testing miss, 38% said people-centered review would receive the fastest investment growth, compared with 24% of organizations that had not been burned.

                                      That produces an apparent paradox: the burned group is most likely to remove people from the release checkpoint and most likely to increase spending on people elsewhere in the process. The strategy appears to be automation with a human backstop — allow agents to move faster, then use reviewers to catch what automated evaluation misses.

                                      The open question is whether that model scales. Agent deployments and automated checks can grow with software volume. Reviewer hours do not fall at the same rate. Enterprises may therefore be replacing a human approval step with a larger downstream review function rather than eliminating human oversight.

                                      The narrow but consequential read

                                      July’s data does not show that enterprise agent evaluation is failing everywhere, nor does it prove automated judges are getting worse. It shows something more precise: confidence rose before the measured failure incidence improved.

                                      At the same time, the infrastructure around evaluation is maturing. More enterprises are adopting specialist tools, integration has become the leading purchase criterion and companies that have already experienced failures are increasing investment in human review. The market recognizes the problem and is spending against it.

                                      But the central reliability result remains stubborn. Nearly half of surveyed enterprises still report that an AI feature cleared internal checks before disappointing a customer. Respondents with that experience place less faith in automation — and move faster toward deployments with no human approval.

                                      The report frames this as an incomplete verification model: a passing pre-deployment score marks the start of monitoring, not the end of it. For most enterprises, the production quality checks and evaluation testing that would close that gap still aren’t in place.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows.

                                      In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior. Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month.

                                      Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers, essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once.

                                      The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much autonomy companies give agents and how well they can verify them. It is increasingly a gap between confidence in the evaluation layer and evidence that the layer is getting better at preventing failures.

                                      The most revealing split appears inside the July data.

                                      Of the enterprises that experienced an AI feature clear testing only to go on to disappoint a customer, 4% placed complete faith in automated checks. Of those that had detected no comparable incident, 24% expressed full confidence — a sixfold difference.

                                      It makes sense: those who experienced test-passing agents failing in live production are, unsurprisingly, more likely to doubt the automated checking process.

                                      Perhaps it makes sense then, that companies geared toward tackling this problem — like automated agent error monitoring and mitigation platform Raindrop.ai — are seeing the market transform wildly from just a few months ago.

                                      «We are seeing the great-decline of evals as we know them,» Raindrop CTO Ben Hylak told VentureBeat in a direct message. «The Fortune 100 are increasingly reducing eval sets and deprioritizing maintenance. As systems grow more complex (MCPs, subagents, etc.) it becomes impossible to fully enumerate the failure cases. Instead, they’re leaning on anomaly and issue detection solutions, both before and after production.»

                                      A directional finding, not a market census

                                      VentureBeat fielded the July wave among 108 people representing companies with workforces of at least 100. This is down from 157 respondents in June.

                                      Of the 108, 69% described themselves as final AI-buying authorities or people who recommend and influence those purchases. The sample skewed toward midsize organizations: 63% worked at companies with 100 to 2,499 employees.

                                      The findings should be read directionally. The survey is self-selected rather than a probability sample, and the burned-vs.-unburned splits cited throughout this piece rest on groups of 41 to 53 respondents, and other cross-tabs in the report range from 40 to 68.

                                      The industry mix also changed: technology and software participation declined nine points, ending at 14%, while the retail and consumer share added four points and ended at 19%.

                                      The report nevertheless identifies four month-to-month changes worth noticing: more respondents professing complete confidence, fewer naming poor real-world alignment, more choosing integration ease as the decisive buying factor, and Braintrust gaining primary-platform share.

                                      Confidence in automated evals improved, but outcomes stayed flat

                                      VentureBeat’s June research identified an enterprise evaluation gap: companies were granting agents more authority faster than they were developing reliable ways to test them.

                                      July preserves the key number from that first wave. Across 265 enterprise responses over the two months, the proportion reporting at least one test-approved system that disappointed customers stayed within a single percentage point: 50% in June and 49% in July.

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      This figure does not mean that 49% of all agent runs fail, or that any particular evaluation product has a 49% failure rate. The survey asks whether an organization experienced at least one customer-facing incident in the previous year after an AI feature passed its internal tests. Companies that deploy far more agents have more opportunities to encounter such an incident.

                                      But that limitation does not make the result less important. An internal evaluation serves as a release gate. If roughly half of surveyed organizations have seen that gate approve a system that later fails in front of customers, a passing score cannot be treated as proof of production reliability.

                                      The cross-tab reinforces the point. Ten of the 41 enterprises with no identified testing miss placed complete faith in automation.

                                      Only two of the 53 previously burned enterprises said the same. Confidence is strongest among respondents with the least evidence that the release gate can fail.

                                      The enterprises that got burned are moving faster toward zero-human deployment

                                      The counterintuitive finding is what companies do after an evaluation miss.

                                      Overall, 67% either let an agent push code or change a system without a person’s approval in certain low-risk cases, or are modifying their pipelines to support that practice during the coming year. That is unchanged from June. In July, 37% already permitted it in limited cases and another 30% were building toward it.

                                      Among enterprises where a test-approved system had disappointed a customer, however, 85% were pursuing that no-approval model, compared with 61% in the group reporting no comparable incident. Only 11% of burned respondents rejected end-to-end deployment automation for the years ahead, versus 24% of unburned respondents.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 chart

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026.

                                      It would be easy to read that as recklessness, but the data supports another plausible explanation: deployment maturity. Organizations running more agents, at higher volume and across more consequential workflows, are more likely both to encounter failures and to have the engineering infrastructure needed for automated deployment.

                                      The survey cannot establish which explanation dominates. It does establish that a customer-visible incident does not appear to stop the move toward autonomy. Respondents with firsthand proof that testing can miss defects are also moving most aggressively to let those tests authorize production changes.

                                      If their per-deployment failure rate remains constant while deployment volume rises, the total incident count could grow even without the percentage of affected companies increasing. The July data does not measure incident volume, so that remains a risk implied by the pattern rather than a measured outcome.

                                      The release gate is automated, but production quality monitoring still lags

                                      Pre-deployment evaluation and production monitoring answer different questions. An evaluation asks whether an agent appears ready to ship. Production monitoring asks what the agent is doing after release and whether its live outputs remain correct.

                                      Most companies in the July sample still emphasize whether the system functions, not whether the answer is correct.

                                      Among the 106 valid responses to this question, 26% used inline quality assertions — automated judges or guardrails checking live traffic for output-quality problems. Another 26% focused on transaction traces such as infrastructure spans, token usage and raw inputs and outputs, while 24% mainly tracked gateway metrics such as latency, errors and cost.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026 monitoring uptime vs. wrong answers

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026

                                      Trace and gateway data can reveal outages, slowdowns and broken requests. They may not flag a fluent, fast and confidently wrong answer. Grouped by the report according to what each architecture actually watches, half of respondents monitored whether an agent was functioning, while just over a quarter automatically monitored whether its production output was correct.

                                      The gap is sharpest among the 40 respondents already permitting no-approval deployment in limited cases. Only 28% of that group automatically checked the meaning and correctness of live answers. In other words, most enterprises that have eliminated a person from at least some release decisions have not installed automated semantic-quality monitoring as the production backstop.

                                      This is the clearest operational lesson in the data. A pre-deployment test suite and infrastructure observability are necessary, but they do not cover the same failure mode. Enterprises need a way to detect bad outputs after the agent begins interacting with real users, data and tools — especially when nobody reviews the deployment decision first.

                                      An independent agent-evaluation market begins to take shape

                                      The vendor data offers a more encouraging sign: enterprises are adding dedicated evaluation tools, and specialist platforms are gaining ground.

                                      OpenAI’s native evals and traces narrowly led as the primary platform at 18%, followed by Confident AI’s DeepEval at 17% and Braintrust at 15%.

                                      Anthropic’s Claude Console and Workbench held 12%, tied with organizations reporting no dedicated evaluation platform. Three options each held 6%: internally built tools, Promptfoo and LangSmith.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 market landscape

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Braintrust’s primary share increased from 8% in June to 15% in July, the biggest gain by one vendor and the one the report flags as statistically significant. DeepEval rose from 12% to 17%. Use of no purpose-built platform declined five points to 12%, although that smaller change does not by itself confirm a trend.

                                      Because many companies use more than one tool, the broader footprints are larger. OpenAI native evaluation appeared somewhere in 31% of stacks, DeepEval in 27%, Braintrust in 22% and Anthropic’s native tooling in 20%. Custom internal tooling reached 14%, while Weights & Biases Weave and open-source Langfuse each reached 11%.

                                      These are adoption figures, not product-performance scores. The survey does not establish that one vendor produces more reliable agents than another. Still, the results point toward evaluation becoming a distinct enterprise software layer rather than a loose collection of internal scripts or a feature used only inside a model provider’s platform.

                                      Purchasing priorities are changing with that market. The proportion choosing integration ease as the decisive factor climbed 12 points to 39%, displacing cost, which fell from 28% to 23%. Evaluation accuracy ranked second at 28%. The combined average for satisfaction, implementation simplicity and economic value was 3.9 out of five.

                                      The move from price toward integration suggests enterprises increasingly want a tool they can install into existing development and monitoring pipelines now. Yet their leading success metric remains evaluation consistency at 38%, followed by fewer failures and regressions at 20%. Buyers are selecting for fit while still judging results on repeatability.

                                      Switching intent also cooled: 56% still expected to add or replace a platform during the coming year, down from 64% in June.

                                      The proportion staying put increased eight points, reaching 44%. Together with specialist adoption, they suggest some buyers are moving from evaluation to implementation.

                                      Human review is becoming the hedge against automated misses

                                      The budget data reveals how enterprises are managing the contradiction between greater autonomy and unreliable evaluation.

                                      People-centered review workflows edged narrowly ahead of production observability as the most frequently cited area for increased investment, 31% to 30%.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 human review

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Automated evaluation pipelines ranked third at 19%, followed by testing for safety and policy compliance at 16%. Only 6% said their reliability and evaluation budget was not increasing.

                                      Among enterprises that had experienced a testing miss, 38% said people-centered review would receive the fastest investment growth, compared with 24% of organizations that had not been burned.

                                      That produces an apparent paradox: the burned group is most likely to remove people from the release checkpoint and most likely to increase spending on people elsewhere in the process. The strategy appears to be automation with a human backstop — allow agents to move faster, then use reviewers to catch what automated evaluation misses.

                                      The open question is whether that model scales. Agent deployments and automated checks can grow with software volume. Reviewer hours do not fall at the same rate. Enterprises may therefore be replacing a human approval step with a larger downstream review function rather than eliminating human oversight.

                                      The narrow but consequential read

                                      July’s data does not show that enterprise agent evaluation is failing everywhere, nor does it prove automated judges are getting worse. It shows something more precise: confidence rose before the measured failure incidence improved.

                                      At the same time, the infrastructure around evaluation is maturing. More enterprises are adopting specialist tools, integration has become the leading purchase criterion and companies that have already experienced failures are increasing investment in human review. The market recognizes the problem and is spending against it.

                                      But the central reliability result remains stubborn. Nearly half of surveyed enterprises still report that an AI feature cleared internal checks before disappointing a customer. Respondents with that experience place less faith in automation — and move faster toward deployments with no human approval.

                                      The report frames this as an incomplete verification model: a passing pre-deployment score marks the start of monitoring, not the end of it. For most enterprises, the production quality checks and evaluation testing that would close that gap still aren’t in place.

                                      Tours Colombia Todo el año Tours Colombia Todo el año Tours Colombia Todo el año

                                      Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows.

                                      In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior. Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month.

                                      Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers, essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once.

                                      The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much autonomy companies give agents and how well they can verify them. It is increasingly a gap between confidence in the evaluation layer and evidence that the layer is getting better at preventing failures.

                                      The most revealing split appears inside the July data.

                                      Of the enterprises that experienced an AI feature clear testing only to go on to disappoint a customer, 4% placed complete faith in automated checks. Of those that had detected no comparable incident, 24% expressed full confidence — a sixfold difference.

                                      It makes sense: those who experienced test-passing agents failing in live production are, unsurprisingly, more likely to doubt the automated checking process.

                                      Perhaps it makes sense then, that companies geared toward tackling this problem — like automated agent error monitoring and mitigation platform Raindrop.ai — are seeing the market transform wildly from just a few months ago.

                                      «We are seeing the great-decline of evals as we know them,» Raindrop CTO Ben Hylak told VentureBeat in a direct message. «The Fortune 100 are increasingly reducing eval sets and deprioritizing maintenance. As systems grow more complex (MCPs, subagents, etc.) it becomes impossible to fully enumerate the failure cases. Instead, they’re leaning on anomaly and issue detection solutions, both before and after production.»

                                      A directional finding, not a market census

                                      VentureBeat fielded the July wave among 108 people representing companies with workforces of at least 100. This is down from 157 respondents in June.

                                      Of the 108, 69% described themselves as final AI-buying authorities or people who recommend and influence those purchases. The sample skewed toward midsize organizations: 63% worked at companies with 100 to 2,499 employees.

                                      The findings should be read directionally. The survey is self-selected rather than a probability sample, and the burned-vs.-unburned splits cited throughout this piece rest on groups of 41 to 53 respondents, and other cross-tabs in the report range from 40 to 68.

                                      The industry mix also changed: technology and software participation declined nine points, ending at 14%, while the retail and consumer share added four points and ended at 19%.

                                      The report nevertheless identifies four month-to-month changes worth noticing: more respondents professing complete confidence, fewer naming poor real-world alignment, more choosing integration ease as the decisive buying factor, and Braintrust gaining primary-platform share.

                                      Confidence in automated evals improved, but outcomes stayed flat

                                      VentureBeat’s June research identified an enterprise evaluation gap: companies were granting agents more authority faster than they were developing reliable ways to test them.

                                      July preserves the key number from that first wave. Across 265 enterprise responses over the two months, the proportion reporting at least one test-approved system that disappointed customers stayed within a single percentage point: 50% in June and 49% in July.

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      This figure does not mean that 49% of all agent runs fail, or that any particular evaluation product has a 49% failure rate. The survey asks whether an organization experienced at least one customer-facing incident in the previous year after an AI feature passed its internal tests. Companies that deploy far more agents have more opportunities to encounter such an incident.

                                      But that limitation does not make the result less important. An internal evaluation serves as a release gate. If roughly half of surveyed organizations have seen that gate approve a system that later fails in front of customers, a passing score cannot be treated as proof of production reliability.

                                      The cross-tab reinforces the point. Ten of the 41 enterprises with no identified testing miss placed complete faith in automation.

                                      Only two of the 53 previously burned enterprises said the same. Confidence is strongest among respondents with the least evidence that the release gate can fail.

                                      The enterprises that got burned are moving faster toward zero-human deployment

                                      The counterintuitive finding is what companies do after an evaluation miss.

                                      Overall, 67% either let an agent push code or change a system without a person’s approval in certain low-risk cases, or are modifying their pipelines to support that practice during the coming year. That is unchanged from June. In July, 37% already permitted it in limited cases and another 30% were building toward it.

                                      Among enterprises where a test-approved system had disappointed a customer, however, 85% were pursuing that no-approval model, compared with 61% in the group reporting no comparable incident. Only 11% of burned respondents rejected end-to-end deployment automation for the years ahead, versus 24% of unburned respondents.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 chart

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026.

                                      It would be easy to read that as recklessness, but the data supports another plausible explanation: deployment maturity. Organizations running more agents, at higher volume and across more consequential workflows, are more likely both to encounter failures and to have the engineering infrastructure needed for automated deployment.

                                      The survey cannot establish which explanation dominates. It does establish that a customer-visible incident does not appear to stop the move toward autonomy. Respondents with firsthand proof that testing can miss defects are also moving most aggressively to let those tests authorize production changes.

                                      If their per-deployment failure rate remains constant while deployment volume rises, the total incident count could grow even without the percentage of affected companies increasing. The July data does not measure incident volume, so that remains a risk implied by the pattern rather than a measured outcome.

                                      The release gate is automated, but production quality monitoring still lags

                                      Pre-deployment evaluation and production monitoring answer different questions. An evaluation asks whether an agent appears ready to ship. Production monitoring asks what the agent is doing after release and whether its live outputs remain correct.

                                      Most companies in the July sample still emphasize whether the system functions, not whether the answer is correct.

                                      Among the 106 valid responses to this question, 26% used inline quality assertions — automated judges or guardrails checking live traffic for output-quality problems. Another 26% focused on transaction traces such as infrastructure spans, token usage and raw inputs and outputs, while 24% mainly tracked gateway metrics such as latency, errors and cost.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026 monitoring uptime vs. wrong answers

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026

                                      Trace and gateway data can reveal outages, slowdowns and broken requests. They may not flag a fluent, fast and confidently wrong answer. Grouped by the report according to what each architecture actually watches, half of respondents monitored whether an agent was functioning, while just over a quarter automatically monitored whether its production output was correct.

                                      The gap is sharpest among the 40 respondents already permitting no-approval deployment in limited cases. Only 28% of that group automatically checked the meaning and correctness of live answers. In other words, most enterprises that have eliminated a person from at least some release decisions have not installed automated semantic-quality monitoring as the production backstop.

                                      This is the clearest operational lesson in the data. A pre-deployment test suite and infrastructure observability are necessary, but they do not cover the same failure mode. Enterprises need a way to detect bad outputs after the agent begins interacting with real users, data and tools — especially when nobody reviews the deployment decision first.

                                      An independent agent-evaluation market begins to take shape

                                      The vendor data offers a more encouraging sign: enterprises are adding dedicated evaluation tools, and specialist platforms are gaining ground.

                                      OpenAI’s native evals and traces narrowly led as the primary platform at 18%, followed by Confident AI’s DeepEval at 17% and Braintrust at 15%.

                                      Anthropic’s Claude Console and Workbench held 12%, tied with organizations reporting no dedicated evaluation platform. Three options each held 6%: internally built tools, Promptfoo and LangSmith.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 market landscape

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Braintrust’s primary share increased from 8% in June to 15% in July, the biggest gain by one vendor and the one the report flags as statistically significant. DeepEval rose from 12% to 17%. Use of no purpose-built platform declined five points to 12%, although that smaller change does not by itself confirm a trend.

                                      Because many companies use more than one tool, the broader footprints are larger. OpenAI native evaluation appeared somewhere in 31% of stacks, DeepEval in 27%, Braintrust in 22% and Anthropic’s native tooling in 20%. Custom internal tooling reached 14%, while Weights & Biases Weave and open-source Langfuse each reached 11%.

                                      These are adoption figures, not product-performance scores. The survey does not establish that one vendor produces more reliable agents than another. Still, the results point toward evaluation becoming a distinct enterprise software layer rather than a loose collection of internal scripts or a feature used only inside a model provider’s platform.

                                      Purchasing priorities are changing with that market. The proportion choosing integration ease as the decisive factor climbed 12 points to 39%, displacing cost, which fell from 28% to 23%. Evaluation accuracy ranked second at 28%. The combined average for satisfaction, implementation simplicity and economic value was 3.9 out of five.

                                      The move from price toward integration suggests enterprises increasingly want a tool they can install into existing development and monitoring pipelines now. Yet their leading success metric remains evaluation consistency at 38%, followed by fewer failures and regressions at 20%. Buyers are selecting for fit while still judging results on repeatability.

                                      Switching intent also cooled: 56% still expected to add or replace a platform during the coming year, down from 64% in June.

                                      The proportion staying put increased eight points, reaching 44%. Together with specialist adoption, they suggest some buyers are moving from evaluation to implementation.

                                      Human review is becoming the hedge against automated misses

                                      The budget data reveals how enterprises are managing the contradiction between greater autonomy and unreliable evaluation.

                                      People-centered review workflows edged narrowly ahead of production observability as the most frequently cited area for increased investment, 31% to 30%.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 human review

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Automated evaluation pipelines ranked third at 19%, followed by testing for safety and policy compliance at 16%. Only 6% said their reliability and evaluation budget was not increasing.

                                      Among enterprises that had experienced a testing miss, 38% said people-centered review would receive the fastest investment growth, compared with 24% of organizations that had not been burned.

                                      That produces an apparent paradox: the burned group is most likely to remove people from the release checkpoint and most likely to increase spending on people elsewhere in the process. The strategy appears to be automation with a human backstop — allow agents to move faster, then use reviewers to catch what automated evaluation misses.

                                      The open question is whether that model scales. Agent deployments and automated checks can grow with software volume. Reviewer hours do not fall at the same rate. Enterprises may therefore be replacing a human approval step with a larger downstream review function rather than eliminating human oversight.

                                      The narrow but consequential read

                                      July’s data does not show that enterprise agent evaluation is failing everywhere, nor does it prove automated judges are getting worse. It shows something more precise: confidence rose before the measured failure incidence improved.

                                      At the same time, the infrastructure around evaluation is maturing. More enterprises are adopting specialist tools, integration has become the leading purchase criterion and companies that have already experienced failures are increasing investment in human review. The market recognizes the problem and is spending against it.

                                      But the central reliability result remains stubborn. Nearly half of surveyed enterprises still report that an AI feature cleared internal checks before disappointing a customer. Respondents with that experience place less faith in automation — and move faster toward deployments with no human approval.

                                      The report frames this as an incomplete verification model: a passing pre-deployment score marks the start of monitoring, not the end of it. For most enterprises, the production quality checks and evaluation testing that would close that gap still aren’t in place.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows.

                                      In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior. Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month.

                                      Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers, essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once.

                                      The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much autonomy companies give agents and how well they can verify them. It is increasingly a gap between confidence in the evaluation layer and evidence that the layer is getting better at preventing failures.

                                      The most revealing split appears inside the July data.

                                      Of the enterprises that experienced an AI feature clear testing only to go on to disappoint a customer, 4% placed complete faith in automated checks. Of those that had detected no comparable incident, 24% expressed full confidence — a sixfold difference.

                                      It makes sense: those who experienced test-passing agents failing in live production are, unsurprisingly, more likely to doubt the automated checking process.

                                      Perhaps it makes sense then, that companies geared toward tackling this problem — like automated agent error monitoring and mitigation platform Raindrop.ai — are seeing the market transform wildly from just a few months ago.

                                      «We are seeing the great-decline of evals as we know them,» Raindrop CTO Ben Hylak told VentureBeat in a direct message. «The Fortune 100 are increasingly reducing eval sets and deprioritizing maintenance. As systems grow more complex (MCPs, subagents, etc.) it becomes impossible to fully enumerate the failure cases. Instead, they’re leaning on anomaly and issue detection solutions, both before and after production.»

                                      A directional finding, not a market census

                                      VentureBeat fielded the July wave among 108 people representing companies with workforces of at least 100. This is down from 157 respondents in June.

                                      Of the 108, 69% described themselves as final AI-buying authorities or people who recommend and influence those purchases. The sample skewed toward midsize organizations: 63% worked at companies with 100 to 2,499 employees.

                                      The findings should be read directionally. The survey is self-selected rather than a probability sample, and the burned-vs.-unburned splits cited throughout this piece rest on groups of 41 to 53 respondents, and other cross-tabs in the report range from 40 to 68.

                                      The industry mix also changed: technology and software participation declined nine points, ending at 14%, while the retail and consumer share added four points and ended at 19%.

                                      The report nevertheless identifies four month-to-month changes worth noticing: more respondents professing complete confidence, fewer naming poor real-world alignment, more choosing integration ease as the decisive buying factor, and Braintrust gaining primary-platform share.

                                      Confidence in automated evals improved, but outcomes stayed flat

                                      VentureBeat’s June research identified an enterprise evaluation gap: companies were granting agents more authority faster than they were developing reliable ways to test them.

                                      July preserves the key number from that first wave. Across 265 enterprise responses over the two months, the proportion reporting at least one test-approved system that disappointed customers stayed within a single percentage point: 50% in June and 49% in July.

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      This figure does not mean that 49% of all agent runs fail, or that any particular evaluation product has a 49% failure rate. The survey asks whether an organization experienced at least one customer-facing incident in the previous year after an AI feature passed its internal tests. Companies that deploy far more agents have more opportunities to encounter such an incident.

                                      But that limitation does not make the result less important. An internal evaluation serves as a release gate. If roughly half of surveyed organizations have seen that gate approve a system that later fails in front of customers, a passing score cannot be treated as proof of production reliability.

                                      The cross-tab reinforces the point. Ten of the 41 enterprises with no identified testing miss placed complete faith in automation.

                                      Only two of the 53 previously burned enterprises said the same. Confidence is strongest among respondents with the least evidence that the release gate can fail.

                                      The enterprises that got burned are moving faster toward zero-human deployment

                                      The counterintuitive finding is what companies do after an evaluation miss.

                                      Overall, 67% either let an agent push code or change a system without a person’s approval in certain low-risk cases, or are modifying their pipelines to support that practice during the coming year. That is unchanged from June. In July, 37% already permitted it in limited cases and another 30% were building toward it.

                                      Among enterprises where a test-approved system had disappointed a customer, however, 85% were pursuing that no-approval model, compared with 61% in the group reporting no comparable incident. Only 11% of burned respondents rejected end-to-end deployment automation for the years ahead, versus 24% of unburned respondents.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 chart

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026.

                                      It would be easy to read that as recklessness, but the data supports another plausible explanation: deployment maturity. Organizations running more agents, at higher volume and across more consequential workflows, are more likely both to encounter failures and to have the engineering infrastructure needed for automated deployment.

                                      The survey cannot establish which explanation dominates. It does establish that a customer-visible incident does not appear to stop the move toward autonomy. Respondents with firsthand proof that testing can miss defects are also moving most aggressively to let those tests authorize production changes.

                                      If their per-deployment failure rate remains constant while deployment volume rises, the total incident count could grow even without the percentage of affected companies increasing. The July data does not measure incident volume, so that remains a risk implied by the pattern rather than a measured outcome.

                                      The release gate is automated, but production quality monitoring still lags

                                      Pre-deployment evaluation and production monitoring answer different questions. An evaluation asks whether an agent appears ready to ship. Production monitoring asks what the agent is doing after release and whether its live outputs remain correct.

                                      Most companies in the July sample still emphasize whether the system functions, not whether the answer is correct.

                                      Among the 106 valid responses to this question, 26% used inline quality assertions — automated judges or guardrails checking live traffic for output-quality problems. Another 26% focused on transaction traces such as infrastructure spans, token usage and raw inputs and outputs, while 24% mainly tracked gateway metrics such as latency, errors and cost.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026 monitoring uptime vs. wrong answers

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026

                                      Trace and gateway data can reveal outages, slowdowns and broken requests. They may not flag a fluent, fast and confidently wrong answer. Grouped by the report according to what each architecture actually watches, half of respondents monitored whether an agent was functioning, while just over a quarter automatically monitored whether its production output was correct.

                                      The gap is sharpest among the 40 respondents already permitting no-approval deployment in limited cases. Only 28% of that group automatically checked the meaning and correctness of live answers. In other words, most enterprises that have eliminated a person from at least some release decisions have not installed automated semantic-quality monitoring as the production backstop.

                                      This is the clearest operational lesson in the data. A pre-deployment test suite and infrastructure observability are necessary, but they do not cover the same failure mode. Enterprises need a way to detect bad outputs after the agent begins interacting with real users, data and tools — especially when nobody reviews the deployment decision first.

                                      An independent agent-evaluation market begins to take shape

                                      The vendor data offers a more encouraging sign: enterprises are adding dedicated evaluation tools, and specialist platforms are gaining ground.

                                      OpenAI’s native evals and traces narrowly led as the primary platform at 18%, followed by Confident AI’s DeepEval at 17% and Braintrust at 15%.

                                      Anthropic’s Claude Console and Workbench held 12%, tied with organizations reporting no dedicated evaluation platform. Three options each held 6%: internally built tools, Promptfoo and LangSmith.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 market landscape

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Braintrust’s primary share increased from 8% in June to 15% in July, the biggest gain by one vendor and the one the report flags as statistically significant. DeepEval rose from 12% to 17%. Use of no purpose-built platform declined five points to 12%, although that smaller change does not by itself confirm a trend.

                                      Because many companies use more than one tool, the broader footprints are larger. OpenAI native evaluation appeared somewhere in 31% of stacks, DeepEval in 27%, Braintrust in 22% and Anthropic’s native tooling in 20%. Custom internal tooling reached 14%, while Weights & Biases Weave and open-source Langfuse each reached 11%.

                                      These are adoption figures, not product-performance scores. The survey does not establish that one vendor produces more reliable agents than another. Still, the results point toward evaluation becoming a distinct enterprise software layer rather than a loose collection of internal scripts or a feature used only inside a model provider’s platform.

                                      Purchasing priorities are changing with that market. The proportion choosing integration ease as the decisive factor climbed 12 points to 39%, displacing cost, which fell from 28% to 23%. Evaluation accuracy ranked second at 28%. The combined average for satisfaction, implementation simplicity and economic value was 3.9 out of five.

                                      The move from price toward integration suggests enterprises increasingly want a tool they can install into existing development and monitoring pipelines now. Yet their leading success metric remains evaluation consistency at 38%, followed by fewer failures and regressions at 20%. Buyers are selecting for fit while still judging results on repeatability.

                                      Switching intent also cooled: 56% still expected to add or replace a platform during the coming year, down from 64% in June.

                                      The proportion staying put increased eight points, reaching 44%. Together with specialist adoption, they suggest some buyers are moving from evaluation to implementation.

                                      Human review is becoming the hedge against automated misses

                                      The budget data reveals how enterprises are managing the contradiction between greater autonomy and unreliable evaluation.

                                      People-centered review workflows edged narrowly ahead of production observability as the most frequently cited area for increased investment, 31% to 30%.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 human review

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Automated evaluation pipelines ranked third at 19%, followed by testing for safety and policy compliance at 16%. Only 6% said their reliability and evaluation budget was not increasing.

                                      Among enterprises that had experienced a testing miss, 38% said people-centered review would receive the fastest investment growth, compared with 24% of organizations that had not been burned.

                                      That produces an apparent paradox: the burned group is most likely to remove people from the release checkpoint and most likely to increase spending on people elsewhere in the process. The strategy appears to be automation with a human backstop — allow agents to move faster, then use reviewers to catch what automated evaluation misses.

                                      The open question is whether that model scales. Agent deployments and automated checks can grow with software volume. Reviewer hours do not fall at the same rate. Enterprises may therefore be replacing a human approval step with a larger downstream review function rather than eliminating human oversight.

                                      The narrow but consequential read

                                      July’s data does not show that enterprise agent evaluation is failing everywhere, nor does it prove automated judges are getting worse. It shows something more precise: confidence rose before the measured failure incidence improved.

                                      At the same time, the infrastructure around evaluation is maturing. More enterprises are adopting specialist tools, integration has become the leading purchase criterion and companies that have already experienced failures are increasing investment in human review. The market recognizes the problem and is spending against it.

                                      But the central reliability result remains stubborn. Nearly half of surveyed enterprises still report that an AI feature cleared internal checks before disappointing a customer. Respondents with that experience place less faith in automation — and move faster toward deployments with no human approval.

                                      The report frames this as an incomplete verification model: a passing pre-deployment score marks the start of monitoring, not the end of it. For most enterprises, the production quality checks and evaluation testing that would close that gap still aren’t in place.

                                      ¡No te pierdas las noticias destacadas!

                                      Suscríbete y recibe las historias más importantes del día.

                                      Al suscribirte aceptas nuestros términos y condiciones y política de privacidad.

                                      Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows.

                                      In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior. Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month.

                                      Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers, essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once.

                                      The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much autonomy companies give agents and how well they can verify them. It is increasingly a gap between confidence in the evaluation layer and evidence that the layer is getting better at preventing failures.

                                      The most revealing split appears inside the July data.

                                      Of the enterprises that experienced an AI feature clear testing only to go on to disappoint a customer, 4% placed complete faith in automated checks. Of those that had detected no comparable incident, 24% expressed full confidence — a sixfold difference.

                                      It makes sense: those who experienced test-passing agents failing in live production are, unsurprisingly, more likely to doubt the automated checking process.

                                      Perhaps it makes sense then, that companies geared toward tackling this problem — like automated agent error monitoring and mitigation platform Raindrop.ai — are seeing the market transform wildly from just a few months ago.

                                      «We are seeing the great-decline of evals as we know them,» Raindrop CTO Ben Hylak told VentureBeat in a direct message. «The Fortune 100 are increasingly reducing eval sets and deprioritizing maintenance. As systems grow more complex (MCPs, subagents, etc.) it becomes impossible to fully enumerate the failure cases. Instead, they’re leaning on anomaly and issue detection solutions, both before and after production.»

                                      A directional finding, not a market census

                                      VentureBeat fielded the July wave among 108 people representing companies with workforces of at least 100. This is down from 157 respondents in June.

                                      Of the 108, 69% described themselves as final AI-buying authorities or people who recommend and influence those purchases. The sample skewed toward midsize organizations: 63% worked at companies with 100 to 2,499 employees.

                                      The findings should be read directionally. The survey is self-selected rather than a probability sample, and the burned-vs.-unburned splits cited throughout this piece rest on groups of 41 to 53 respondents, and other cross-tabs in the report range from 40 to 68.

                                      The industry mix also changed: technology and software participation declined nine points, ending at 14%, while the retail and consumer share added four points and ended at 19%.

                                      The report nevertheless identifies four month-to-month changes worth noticing: more respondents professing complete confidence, fewer naming poor real-world alignment, more choosing integration ease as the decisive buying factor, and Braintrust gaining primary-platform share.

                                      Confidence in automated evals improved, but outcomes stayed flat

                                      VentureBeat’s June research identified an enterprise evaluation gap: companies were granting agents more authority faster than they were developing reliable ways to test them.

                                      July preserves the key number from that first wave. Across 265 enterprise responses over the two months, the proportion reporting at least one test-approved system that disappointed customers stayed within a single percentage point: 50% in June and 49% in July.

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      This figure does not mean that 49% of all agent runs fail, or that any particular evaluation product has a 49% failure rate. The survey asks whether an organization experienced at least one customer-facing incident in the previous year after an AI feature passed its internal tests. Companies that deploy far more agents have more opportunities to encounter such an incident.

                                      But that limitation does not make the result less important. An internal evaluation serves as a release gate. If roughly half of surveyed organizations have seen that gate approve a system that later fails in front of customers, a passing score cannot be treated as proof of production reliability.

                                      The cross-tab reinforces the point. Ten of the 41 enterprises with no identified testing miss placed complete faith in automation.

                                      Only two of the 53 previously burned enterprises said the same. Confidence is strongest among respondents with the least evidence that the release gate can fail.

                                      The enterprises that got burned are moving faster toward zero-human deployment

                                      The counterintuitive finding is what companies do after an evaluation miss.

                                      Overall, 67% either let an agent push code or change a system without a person’s approval in certain low-risk cases, or are modifying their pipelines to support that practice during the coming year. That is unchanged from June. In July, 37% already permitted it in limited cases and another 30% were building toward it.

                                      Among enterprises where a test-approved system had disappointed a customer, however, 85% were pursuing that no-approval model, compared with 61% in the group reporting no comparable incident. Only 11% of burned respondents rejected end-to-end deployment automation for the years ahead, versus 24% of unburned respondents.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 chart

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026.

                                      It would be easy to read that as recklessness, but the data supports another plausible explanation: deployment maturity. Organizations running more agents, at higher volume and across more consequential workflows, are more likely both to encounter failures and to have the engineering infrastructure needed for automated deployment.

                                      The survey cannot establish which explanation dominates. It does establish that a customer-visible incident does not appear to stop the move toward autonomy. Respondents with firsthand proof that testing can miss defects are also moving most aggressively to let those tests authorize production changes.

                                      If their per-deployment failure rate remains constant while deployment volume rises, the total incident count could grow even without the percentage of affected companies increasing. The July data does not measure incident volume, so that remains a risk implied by the pattern rather than a measured outcome.

                                      The release gate is automated, but production quality monitoring still lags

                                      Pre-deployment evaluation and production monitoring answer different questions. An evaluation asks whether an agent appears ready to ship. Production monitoring asks what the agent is doing after release and whether its live outputs remain correct.

                                      Most companies in the July sample still emphasize whether the system functions, not whether the answer is correct.

                                      Among the 106 valid responses to this question, 26% used inline quality assertions — automated judges or guardrails checking live traffic for output-quality problems. Another 26% focused on transaction traces such as infrastructure spans, token usage and raw inputs and outputs, while 24% mainly tracked gateway metrics such as latency, errors and cost.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026 monitoring uptime vs. wrong answers

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026

                                      Trace and gateway data can reveal outages, slowdowns and broken requests. They may not flag a fluent, fast and confidently wrong answer. Grouped by the report according to what each architecture actually watches, half of respondents monitored whether an agent was functioning, while just over a quarter automatically monitored whether its production output was correct.

                                      The gap is sharpest among the 40 respondents already permitting no-approval deployment in limited cases. Only 28% of that group automatically checked the meaning and correctness of live answers. In other words, most enterprises that have eliminated a person from at least some release decisions have not installed automated semantic-quality monitoring as the production backstop.

                                      This is the clearest operational lesson in the data. A pre-deployment test suite and infrastructure observability are necessary, but they do not cover the same failure mode. Enterprises need a way to detect bad outputs after the agent begins interacting with real users, data and tools — especially when nobody reviews the deployment decision first.

                                      An independent agent-evaluation market begins to take shape

                                      The vendor data offers a more encouraging sign: enterprises are adding dedicated evaluation tools, and specialist platforms are gaining ground.

                                      OpenAI’s native evals and traces narrowly led as the primary platform at 18%, followed by Confident AI’s DeepEval at 17% and Braintrust at 15%.

                                      Anthropic’s Claude Console and Workbench held 12%, tied with organizations reporting no dedicated evaluation platform. Three options each held 6%: internally built tools, Promptfoo and LangSmith.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 market landscape

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Braintrust’s primary share increased from 8% in June to 15% in July, the biggest gain by one vendor and the one the report flags as statistically significant. DeepEval rose from 12% to 17%. Use of no purpose-built platform declined five points to 12%, although that smaller change does not by itself confirm a trend.

                                      Because many companies use more than one tool, the broader footprints are larger. OpenAI native evaluation appeared somewhere in 31% of stacks, DeepEval in 27%, Braintrust in 22% and Anthropic’s native tooling in 20%. Custom internal tooling reached 14%, while Weights & Biases Weave and open-source Langfuse each reached 11%.

                                      These are adoption figures, not product-performance scores. The survey does not establish that one vendor produces more reliable agents than another. Still, the results point toward evaluation becoming a distinct enterprise software layer rather than a loose collection of internal scripts or a feature used only inside a model provider’s platform.

                                      Purchasing priorities are changing with that market. The proportion choosing integration ease as the decisive factor climbed 12 points to 39%, displacing cost, which fell from 28% to 23%. Evaluation accuracy ranked second at 28%. The combined average for satisfaction, implementation simplicity and economic value was 3.9 out of five.

                                      The move from price toward integration suggests enterprises increasingly want a tool they can install into existing development and monitoring pipelines now. Yet their leading success metric remains evaluation consistency at 38%, followed by fewer failures and regressions at 20%. Buyers are selecting for fit while still judging results on repeatability.

                                      Switching intent also cooled: 56% still expected to add or replace a platform during the coming year, down from 64% in June.

                                      The proportion staying put increased eight points, reaching 44%. Together with specialist adoption, they suggest some buyers are moving from evaluation to implementation.

                                      Human review is becoming the hedge against automated misses

                                      The budget data reveals how enterprises are managing the contradiction between greater autonomy and unreliable evaluation.

                                      People-centered review workflows edged narrowly ahead of production observability as the most frequently cited area for increased investment, 31% to 30%.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 human review

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Automated evaluation pipelines ranked third at 19%, followed by testing for safety and policy compliance at 16%. Only 6% said their reliability and evaluation budget was not increasing.

                                      Among enterprises that had experienced a testing miss, 38% said people-centered review would receive the fastest investment growth, compared with 24% of organizations that had not been burned.

                                      That produces an apparent paradox: the burned group is most likely to remove people from the release checkpoint and most likely to increase spending on people elsewhere in the process. The strategy appears to be automation with a human backstop — allow agents to move faster, then use reviewers to catch what automated evaluation misses.

                                      The open question is whether that model scales. Agent deployments and automated checks can grow with software volume. Reviewer hours do not fall at the same rate. Enterprises may therefore be replacing a human approval step with a larger downstream review function rather than eliminating human oversight.

                                      The narrow but consequential read

                                      July’s data does not show that enterprise agent evaluation is failing everywhere, nor does it prove automated judges are getting worse. It shows something more precise: confidence rose before the measured failure incidence improved.

                                      At the same time, the infrastructure around evaluation is maturing. More enterprises are adopting specialist tools, integration has become the leading purchase criterion and companies that have already experienced failures are increasing investment in human review. The market recognizes the problem and is spending against it.

                                      But the central reliability result remains stubborn. Nearly half of surveyed enterprises still report that an AI feature cleared internal checks before disappointing a customer. Respondents with that experience place less faith in automation — and move faster toward deployments with no human approval.

                                      The report frames this as an incomplete verification model: a passing pre-deployment score marks the start of monitoring, not the end of it. For most enterprises, the production quality checks and evaluation testing that would close that gap still aren’t in place.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows.

                                      In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior. Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month.

                                      Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers, essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once.

                                      The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much autonomy companies give agents and how well they can verify them. It is increasingly a gap between confidence in the evaluation layer and evidence that the layer is getting better at preventing failures.

                                      The most revealing split appears inside the July data.

                                      Of the enterprises that experienced an AI feature clear testing only to go on to disappoint a customer, 4% placed complete faith in automated checks. Of those that had detected no comparable incident, 24% expressed full confidence — a sixfold difference.

                                      It makes sense: those who experienced test-passing agents failing in live production are, unsurprisingly, more likely to doubt the automated checking process.

                                      Perhaps it makes sense then, that companies geared toward tackling this problem — like automated agent error monitoring and mitigation platform Raindrop.ai — are seeing the market transform wildly from just a few months ago.

                                      «We are seeing the great-decline of evals as we know them,» Raindrop CTO Ben Hylak told VentureBeat in a direct message. «The Fortune 100 are increasingly reducing eval sets and deprioritizing maintenance. As systems grow more complex (MCPs, subagents, etc.) it becomes impossible to fully enumerate the failure cases. Instead, they’re leaning on anomaly and issue detection solutions, both before and after production.»

                                      A directional finding, not a market census

                                      VentureBeat fielded the July wave among 108 people representing companies with workforces of at least 100. This is down from 157 respondents in June.

                                      Of the 108, 69% described themselves as final AI-buying authorities or people who recommend and influence those purchases. The sample skewed toward midsize organizations: 63% worked at companies with 100 to 2,499 employees.

                                      The findings should be read directionally. The survey is self-selected rather than a probability sample, and the burned-vs.-unburned splits cited throughout this piece rest on groups of 41 to 53 respondents, and other cross-tabs in the report range from 40 to 68.

                                      The industry mix also changed: technology and software participation declined nine points, ending at 14%, while the retail and consumer share added four points and ended at 19%.

                                      The report nevertheless identifies four month-to-month changes worth noticing: more respondents professing complete confidence, fewer naming poor real-world alignment, more choosing integration ease as the decisive buying factor, and Braintrust gaining primary-platform share.

                                      Confidence in automated evals improved, but outcomes stayed flat

                                      VentureBeat’s June research identified an enterprise evaluation gap: companies were granting agents more authority faster than they were developing reliable ways to test them.

                                      July preserves the key number from that first wave. Across 265 enterprise responses over the two months, the proportion reporting at least one test-approved system that disappointed customers stayed within a single percentage point: 50% in June and 49% in July.

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      This figure does not mean that 49% of all agent runs fail, or that any particular evaluation product has a 49% failure rate. The survey asks whether an organization experienced at least one customer-facing incident in the previous year after an AI feature passed its internal tests. Companies that deploy far more agents have more opportunities to encounter such an incident.

                                      But that limitation does not make the result less important. An internal evaluation serves as a release gate. If roughly half of surveyed organizations have seen that gate approve a system that later fails in front of customers, a passing score cannot be treated as proof of production reliability.

                                      The cross-tab reinforces the point. Ten of the 41 enterprises with no identified testing miss placed complete faith in automation.

                                      Only two of the 53 previously burned enterprises said the same. Confidence is strongest among respondents with the least evidence that the release gate can fail.

                                      The enterprises that got burned are moving faster toward zero-human deployment

                                      The counterintuitive finding is what companies do after an evaluation miss.

                                      Overall, 67% either let an agent push code or change a system without a person’s approval in certain low-risk cases, or are modifying their pipelines to support that practice during the coming year. That is unchanged from June. In July, 37% already permitted it in limited cases and another 30% were building toward it.

                                      Among enterprises where a test-approved system had disappointed a customer, however, 85% were pursuing that no-approval model, compared with 61% in the group reporting no comparable incident. Only 11% of burned respondents rejected end-to-end deployment automation for the years ahead, versus 24% of unburned respondents.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 chart

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026.

                                      It would be easy to read that as recklessness, but the data supports another plausible explanation: deployment maturity. Organizations running more agents, at higher volume and across more consequential workflows, are more likely both to encounter failures and to have the engineering infrastructure needed for automated deployment.

                                      The survey cannot establish which explanation dominates. It does establish that a customer-visible incident does not appear to stop the move toward autonomy. Respondents with firsthand proof that testing can miss defects are also moving most aggressively to let those tests authorize production changes.

                                      If their per-deployment failure rate remains constant while deployment volume rises, the total incident count could grow even without the percentage of affected companies increasing. The July data does not measure incident volume, so that remains a risk implied by the pattern rather than a measured outcome.

                                      The release gate is automated, but production quality monitoring still lags

                                      Pre-deployment evaluation and production monitoring answer different questions. An evaluation asks whether an agent appears ready to ship. Production monitoring asks what the agent is doing after release and whether its live outputs remain correct.

                                      Most companies in the July sample still emphasize whether the system functions, not whether the answer is correct.

                                      Among the 106 valid responses to this question, 26% used inline quality assertions — automated judges or guardrails checking live traffic for output-quality problems. Another 26% focused on transaction traces such as infrastructure spans, token usage and raw inputs and outputs, while 24% mainly tracked gateway metrics such as latency, errors and cost.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026 monitoring uptime vs. wrong answers

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026

                                      Trace and gateway data can reveal outages, slowdowns and broken requests. They may not flag a fluent, fast and confidently wrong answer. Grouped by the report according to what each architecture actually watches, half of respondents monitored whether an agent was functioning, while just over a quarter automatically monitored whether its production output was correct.

                                      The gap is sharpest among the 40 respondents already permitting no-approval deployment in limited cases. Only 28% of that group automatically checked the meaning and correctness of live answers. In other words, most enterprises that have eliminated a person from at least some release decisions have not installed automated semantic-quality monitoring as the production backstop.

                                      This is the clearest operational lesson in the data. A pre-deployment test suite and infrastructure observability are necessary, but they do not cover the same failure mode. Enterprises need a way to detect bad outputs after the agent begins interacting with real users, data and tools — especially when nobody reviews the deployment decision first.

                                      An independent agent-evaluation market begins to take shape

                                      The vendor data offers a more encouraging sign: enterprises are adding dedicated evaluation tools, and specialist platforms are gaining ground.

                                      OpenAI’s native evals and traces narrowly led as the primary platform at 18%, followed by Confident AI’s DeepEval at 17% and Braintrust at 15%.

                                      Anthropic’s Claude Console and Workbench held 12%, tied with organizations reporting no dedicated evaluation platform. Three options each held 6%: internally built tools, Promptfoo and LangSmith.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 market landscape

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Braintrust’s primary share increased from 8% in June to 15% in July, the biggest gain by one vendor and the one the report flags as statistically significant. DeepEval rose from 12% to 17%. Use of no purpose-built platform declined five points to 12%, although that smaller change does not by itself confirm a trend.

                                      Because many companies use more than one tool, the broader footprints are larger. OpenAI native evaluation appeared somewhere in 31% of stacks, DeepEval in 27%, Braintrust in 22% and Anthropic’s native tooling in 20%. Custom internal tooling reached 14%, while Weights & Biases Weave and open-source Langfuse each reached 11%.

                                      These are adoption figures, not product-performance scores. The survey does not establish that one vendor produces more reliable agents than another. Still, the results point toward evaluation becoming a distinct enterprise software layer rather than a loose collection of internal scripts or a feature used only inside a model provider’s platform.

                                      Purchasing priorities are changing with that market. The proportion choosing integration ease as the decisive factor climbed 12 points to 39%, displacing cost, which fell from 28% to 23%. Evaluation accuracy ranked second at 28%. The combined average for satisfaction, implementation simplicity and economic value was 3.9 out of five.

                                      The move from price toward integration suggests enterprises increasingly want a tool they can install into existing development and monitoring pipelines now. Yet their leading success metric remains evaluation consistency at 38%, followed by fewer failures and regressions at 20%. Buyers are selecting for fit while still judging results on repeatability.

                                      Switching intent also cooled: 56% still expected to add or replace a platform during the coming year, down from 64% in June.

                                      The proportion staying put increased eight points, reaching 44%. Together with specialist adoption, they suggest some buyers are moving from evaluation to implementation.

                                      Human review is becoming the hedge against automated misses

                                      The budget data reveals how enterprises are managing the contradiction between greater autonomy and unreliable evaluation.

                                      People-centered review workflows edged narrowly ahead of production observability as the most frequently cited area for increased investment, 31% to 30%.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 human review

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Automated evaluation pipelines ranked third at 19%, followed by testing for safety and policy compliance at 16%. Only 6% said their reliability and evaluation budget was not increasing.

                                      Among enterprises that had experienced a testing miss, 38% said people-centered review would receive the fastest investment growth, compared with 24% of organizations that had not been burned.

                                      That produces an apparent paradox: the burned group is most likely to remove people from the release checkpoint and most likely to increase spending on people elsewhere in the process. The strategy appears to be automation with a human backstop — allow agents to move faster, then use reviewers to catch what automated evaluation misses.

                                      The open question is whether that model scales. Agent deployments and automated checks can grow with software volume. Reviewer hours do not fall at the same rate. Enterprises may therefore be replacing a human approval step with a larger downstream review function rather than eliminating human oversight.

                                      The narrow but consequential read

                                      July’s data does not show that enterprise agent evaluation is failing everywhere, nor does it prove automated judges are getting worse. It shows something more precise: confidence rose before the measured failure incidence improved.

                                      At the same time, the infrastructure around evaluation is maturing. More enterprises are adopting specialist tools, integration has become the leading purchase criterion and companies that have already experienced failures are increasing investment in human review. The market recognizes the problem and is spending against it.

                                      But the central reliability result remains stubborn. Nearly half of surveyed enterprises still report that an AI feature cleared internal checks before disappointing a customer. Respondents with that experience place less faith in automation — and move faster toward deployments with no human approval.

                                      The report frames this as an incomplete verification model: a passing pre-deployment score marks the start of monitoring, not the end of it. For most enterprises, the production quality checks and evaluation testing that would close that gap still aren’t in place.

                                      Tours Colombia Todo el año Tours Colombia Todo el año Tours Colombia Todo el año

                                      Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows.

                                      In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior. Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month.

                                      Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers, essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once.

                                      The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much autonomy companies give agents and how well they can verify them. It is increasingly a gap between confidence in the evaluation layer and evidence that the layer is getting better at preventing failures.

                                      The most revealing split appears inside the July data.

                                      Of the enterprises that experienced an AI feature clear testing only to go on to disappoint a customer, 4% placed complete faith in automated checks. Of those that had detected no comparable incident, 24% expressed full confidence — a sixfold difference.

                                      It makes sense: those who experienced test-passing agents failing in live production are, unsurprisingly, more likely to doubt the automated checking process.

                                      Perhaps it makes sense then, that companies geared toward tackling this problem — like automated agent error monitoring and mitigation platform Raindrop.ai — are seeing the market transform wildly from just a few months ago.

                                      «We are seeing the great-decline of evals as we know them,» Raindrop CTO Ben Hylak told VentureBeat in a direct message. «The Fortune 100 are increasingly reducing eval sets and deprioritizing maintenance. As systems grow more complex (MCPs, subagents, etc.) it becomes impossible to fully enumerate the failure cases. Instead, they’re leaning on anomaly and issue detection solutions, both before and after production.»

                                      A directional finding, not a market census

                                      VentureBeat fielded the July wave among 108 people representing companies with workforces of at least 100. This is down from 157 respondents in June.

                                      Of the 108, 69% described themselves as final AI-buying authorities or people who recommend and influence those purchases. The sample skewed toward midsize organizations: 63% worked at companies with 100 to 2,499 employees.

                                      The findings should be read directionally. The survey is self-selected rather than a probability sample, and the burned-vs.-unburned splits cited throughout this piece rest on groups of 41 to 53 respondents, and other cross-tabs in the report range from 40 to 68.

                                      The industry mix also changed: technology and software participation declined nine points, ending at 14%, while the retail and consumer share added four points and ended at 19%.

                                      The report nevertheless identifies four month-to-month changes worth noticing: more respondents professing complete confidence, fewer naming poor real-world alignment, more choosing integration ease as the decisive buying factor, and Braintrust gaining primary-platform share.

                                      Confidence in automated evals improved, but outcomes stayed flat

                                      VentureBeat’s June research identified an enterprise evaluation gap: companies were granting agents more authority faster than they were developing reliable ways to test them.

                                      July preserves the key number from that first wave. Across 265 enterprise responses over the two months, the proportion reporting at least one test-approved system that disappointed customers stayed within a single percentage point: 50% in June and 49% in July.

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      This figure does not mean that 49% of all agent runs fail, or that any particular evaluation product has a 49% failure rate. The survey asks whether an organization experienced at least one customer-facing incident in the previous year after an AI feature passed its internal tests. Companies that deploy far more agents have more opportunities to encounter such an incident.

                                      But that limitation does not make the result less important. An internal evaluation serves as a release gate. If roughly half of surveyed organizations have seen that gate approve a system that later fails in front of customers, a passing score cannot be treated as proof of production reliability.

                                      The cross-tab reinforces the point. Ten of the 41 enterprises with no identified testing miss placed complete faith in automation.

                                      Only two of the 53 previously burned enterprises said the same. Confidence is strongest among respondents with the least evidence that the release gate can fail.

                                      The enterprises that got burned are moving faster toward zero-human deployment

                                      The counterintuitive finding is what companies do after an evaluation miss.

                                      Overall, 67% either let an agent push code or change a system without a person’s approval in certain low-risk cases, or are modifying their pipelines to support that practice during the coming year. That is unchanged from June. In July, 37% already permitted it in limited cases and another 30% were building toward it.

                                      Among enterprises where a test-approved system had disappointed a customer, however, 85% were pursuing that no-approval model, compared with 61% in the group reporting no comparable incident. Only 11% of burned respondents rejected end-to-end deployment automation for the years ahead, versus 24% of unburned respondents.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 chart

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026.

                                      It would be easy to read that as recklessness, but the data supports another plausible explanation: deployment maturity. Organizations running more agents, at higher volume and across more consequential workflows, are more likely both to encounter failures and to have the engineering infrastructure needed for automated deployment.

                                      The survey cannot establish which explanation dominates. It does establish that a customer-visible incident does not appear to stop the move toward autonomy. Respondents with firsthand proof that testing can miss defects are also moving most aggressively to let those tests authorize production changes.

                                      If their per-deployment failure rate remains constant while deployment volume rises, the total incident count could grow even without the percentage of affected companies increasing. The July data does not measure incident volume, so that remains a risk implied by the pattern rather than a measured outcome.

                                      The release gate is automated, but production quality monitoring still lags

                                      Pre-deployment evaluation and production monitoring answer different questions. An evaluation asks whether an agent appears ready to ship. Production monitoring asks what the agent is doing after release and whether its live outputs remain correct.

                                      Most companies in the July sample still emphasize whether the system functions, not whether the answer is correct.

                                      Among the 106 valid responses to this question, 26% used inline quality assertions — automated judges or guardrails checking live traffic for output-quality problems. Another 26% focused on transaction traces such as infrastructure spans, token usage and raw inputs and outputs, while 24% mainly tracked gateway metrics such as latency, errors and cost.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026 monitoring uptime vs. wrong answers

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026

                                      Trace and gateway data can reveal outages, slowdowns and broken requests. They may not flag a fluent, fast and confidently wrong answer. Grouped by the report according to what each architecture actually watches, half of respondents monitored whether an agent was functioning, while just over a quarter automatically monitored whether its production output was correct.

                                      The gap is sharpest among the 40 respondents already permitting no-approval deployment in limited cases. Only 28% of that group automatically checked the meaning and correctness of live answers. In other words, most enterprises that have eliminated a person from at least some release decisions have not installed automated semantic-quality monitoring as the production backstop.

                                      This is the clearest operational lesson in the data. A pre-deployment test suite and infrastructure observability are necessary, but they do not cover the same failure mode. Enterprises need a way to detect bad outputs after the agent begins interacting with real users, data and tools — especially when nobody reviews the deployment decision first.

                                      An independent agent-evaluation market begins to take shape

                                      The vendor data offers a more encouraging sign: enterprises are adding dedicated evaluation tools, and specialist platforms are gaining ground.

                                      OpenAI’s native evals and traces narrowly led as the primary platform at 18%, followed by Confident AI’s DeepEval at 17% and Braintrust at 15%.

                                      Anthropic’s Claude Console and Workbench held 12%, tied with organizations reporting no dedicated evaluation platform. Three options each held 6%: internally built tools, Promptfoo and LangSmith.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 market landscape

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Braintrust’s primary share increased from 8% in June to 15% in July, the biggest gain by one vendor and the one the report flags as statistically significant. DeepEval rose from 12% to 17%. Use of no purpose-built platform declined five points to 12%, although that smaller change does not by itself confirm a trend.

                                      Because many companies use more than one tool, the broader footprints are larger. OpenAI native evaluation appeared somewhere in 31% of stacks, DeepEval in 27%, Braintrust in 22% and Anthropic’s native tooling in 20%. Custom internal tooling reached 14%, while Weights & Biases Weave and open-source Langfuse each reached 11%.

                                      These are adoption figures, not product-performance scores. The survey does not establish that one vendor produces more reliable agents than another. Still, the results point toward evaluation becoming a distinct enterprise software layer rather than a loose collection of internal scripts or a feature used only inside a model provider’s platform.

                                      Purchasing priorities are changing with that market. The proportion choosing integration ease as the decisive factor climbed 12 points to 39%, displacing cost, which fell from 28% to 23%. Evaluation accuracy ranked second at 28%. The combined average for satisfaction, implementation simplicity and economic value was 3.9 out of five.

                                      The move from price toward integration suggests enterprises increasingly want a tool they can install into existing development and monitoring pipelines now. Yet their leading success metric remains evaluation consistency at 38%, followed by fewer failures and regressions at 20%. Buyers are selecting for fit while still judging results on repeatability.

                                      Switching intent also cooled: 56% still expected to add or replace a platform during the coming year, down from 64% in June.

                                      The proportion staying put increased eight points, reaching 44%. Together with specialist adoption, they suggest some buyers are moving from evaluation to implementation.

                                      Human review is becoming the hedge against automated misses

                                      The budget data reveals how enterprises are managing the contradiction between greater autonomy and unreliable evaluation.

                                      People-centered review workflows edged narrowly ahead of production observability as the most frequently cited area for increased investment, 31% to 30%.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 human review

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Automated evaluation pipelines ranked third at 19%, followed by testing for safety and policy compliance at 16%. Only 6% said their reliability and evaluation budget was not increasing.

                                      Among enterprises that had experienced a testing miss, 38% said people-centered review would receive the fastest investment growth, compared with 24% of organizations that had not been burned.

                                      That produces an apparent paradox: the burned group is most likely to remove people from the release checkpoint and most likely to increase spending on people elsewhere in the process. The strategy appears to be automation with a human backstop — allow agents to move faster, then use reviewers to catch what automated evaluation misses.

                                      The open question is whether that model scales. Agent deployments and automated checks can grow with software volume. Reviewer hours do not fall at the same rate. Enterprises may therefore be replacing a human approval step with a larger downstream review function rather than eliminating human oversight.

                                      The narrow but consequential read

                                      July’s data does not show that enterprise agent evaluation is failing everywhere, nor does it prove automated judges are getting worse. It shows something more precise: confidence rose before the measured failure incidence improved.

                                      At the same time, the infrastructure around evaluation is maturing. More enterprises are adopting specialist tools, integration has become the leading purchase criterion and companies that have already experienced failures are increasing investment in human review. The market recognizes the problem and is spending against it.

                                      But the central reliability result remains stubborn. Nearly half of surveyed enterprises still report that an AI feature cleared internal checks before disappointing a customer. Respondents with that experience place less faith in automation — and move faster toward deployments with no human approval.

                                      The report frames this as an incomplete verification model: a passing pre-deployment score marks the start of monitoring, not the end of it. For most enterprises, the production quality checks and evaluation testing that would close that gap still aren’t in place.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows.

                                      In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior. Meanwhile, survey respondents citing poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%, month over month.

                                      Yet, 49% of survey respondents said that an AI agent or LLM-powered feature that had cleared company testing subsequently created a problem visible to customers, essentially unchanged from 50% in June. And nearly a quarter, 24%, said this troubling outcome had occurred more than once.

                                      The latest findings from VentureBeat Intelligence uncovered a more troubling phase of the enterprise agent rollout: the gap is no longer only between how much autonomy companies give agents and how well they can verify them. It is increasingly a gap between confidence in the evaluation layer and evidence that the layer is getting better at preventing failures.

                                      The most revealing split appears inside the July data.

                                      Of the enterprises that experienced an AI feature clear testing only to go on to disappoint a customer, 4% placed complete faith in automated checks. Of those that had detected no comparable incident, 24% expressed full confidence — a sixfold difference.

                                      It makes sense: those who experienced test-passing agents failing in live production are, unsurprisingly, more likely to doubt the automated checking process.

                                      Perhaps it makes sense then, that companies geared toward tackling this problem — like automated agent error monitoring and mitigation platform Raindrop.ai — are seeing the market transform wildly from just a few months ago.

                                      «We are seeing the great-decline of evals as we know them,» Raindrop CTO Ben Hylak told VentureBeat in a direct message. «The Fortune 100 are increasingly reducing eval sets and deprioritizing maintenance. As systems grow more complex (MCPs, subagents, etc.) it becomes impossible to fully enumerate the failure cases. Instead, they’re leaning on anomaly and issue detection solutions, both before and after production.»

                                      A directional finding, not a market census

                                      VentureBeat fielded the July wave among 108 people representing companies with workforces of at least 100. This is down from 157 respondents in June.

                                      Of the 108, 69% described themselves as final AI-buying authorities or people who recommend and influence those purchases. The sample skewed toward midsize organizations: 63% worked at companies with 100 to 2,499 employees.

                                      The findings should be read directionally. The survey is self-selected rather than a probability sample, and the burned-vs.-unburned splits cited throughout this piece rest on groups of 41 to 53 respondents, and other cross-tabs in the report range from 40 to 68.

                                      The industry mix also changed: technology and software participation declined nine points, ending at 14%, while the retail and consumer share added four points and ended at 19%.

                                      The report nevertheless identifies four month-to-month changes worth noticing: more respondents professing complete confidence, fewer naming poor real-world alignment, more choosing integration ease as the decisive buying factor, and Braintrust gaining primary-platform share.

                                      Confidence in automated evals improved, but outcomes stayed flat

                                      VentureBeat’s June research identified an enterprise evaluation gap: companies were granting agents more authority faster than they were developing reliable ways to test them.

                                      July preserves the key number from that first wave. Across 265 enterprise responses over the two months, the proportion reporting at least one test-approved system that disappointed customers stayed within a single percentage point: 50% in June and 49% in July.

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      This figure does not mean that 49% of all agent runs fail, or that any particular evaluation product has a 49% failure rate. The survey asks whether an organization experienced at least one customer-facing incident in the previous year after an AI feature passed its internal tests. Companies that deploy far more agents have more opportunities to encounter such an incident.

                                      But that limitation does not make the result less important. An internal evaluation serves as a release gate. If roughly half of surveyed organizations have seen that gate approve a system that later fails in front of customers, a passing score cannot be treated as proof of production reliability.

                                      The cross-tab reinforces the point. Ten of the 41 enterprises with no identified testing miss placed complete faith in automation.

                                      Only two of the 53 previously burned enterprises said the same. Confidence is strongest among respondents with the least evidence that the release gate can fail.

                                      The enterprises that got burned are moving faster toward zero-human deployment

                                      The counterintuitive finding is what companies do after an evaluation miss.

                                      Overall, 67% either let an agent push code or change a system without a person’s approval in certain low-risk cases, or are modifying their pipelines to support that practice during the coming year. That is unchanged from June. In July, 37% already permitted it in limited cases and another 30% were building toward it.

                                      Among enterprises where a test-approved system had disappointed a customer, however, 85% were pursuing that no-approval model, compared with 61% in the group reporting no comparable incident. Only 11% of burned respondents rejected end-to-end deployment automation for the years ahead, versus 24% of unburned respondents.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 chart

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026.

                                      It would be easy to read that as recklessness, but the data supports another plausible explanation: deployment maturity. Organizations running more agents, at higher volume and across more consequential workflows, are more likely both to encounter failures and to have the engineering infrastructure needed for automated deployment.

                                      The survey cannot establish which explanation dominates. It does establish that a customer-visible incident does not appear to stop the move toward autonomy. Respondents with firsthand proof that testing can miss defects are also moving most aggressively to let those tests authorize production changes.

                                      If their per-deployment failure rate remains constant while deployment volume rises, the total incident count could grow even without the percentage of affected companies increasing. The July data does not measure incident volume, so that remains a risk implied by the pattern rather than a measured outcome.

                                      The release gate is automated, but production quality monitoring still lags

                                      Pre-deployment evaluation and production monitoring answer different questions. An evaluation asks whether an agent appears ready to ship. Production monitoring asks what the agent is doing after release and whether its live outputs remain correct.

                                      Most companies in the July sample still emphasize whether the system functions, not whether the answer is correct.

                                      Among the 106 valid responses to this question, 26% used inline quality assertions — automated judges or guardrails checking live traffic for output-quality problems. Another 26% focused on transaction traces such as infrastructure spans, token usage and raw inputs and outputs, while 24% mainly tracked gateway metrics such as latency, errors and cost.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026 monitoring uptime vs. wrong answers

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Report Pulse Tracker July 2026

                                      Trace and gateway data can reveal outages, slowdowns and broken requests. They may not flag a fluent, fast and confidently wrong answer. Grouped by the report according to what each architecture actually watches, half of respondents monitored whether an agent was functioning, while just over a quarter automatically monitored whether its production output was correct.

                                      The gap is sharpest among the 40 respondents already permitting no-approval deployment in limited cases. Only 28% of that group automatically checked the meaning and correctness of live answers. In other words, most enterprises that have eliminated a person from at least some release decisions have not installed automated semantic-quality monitoring as the production backstop.

                                      This is the clearest operational lesson in the data. A pre-deployment test suite and infrastructure observability are necessary, but they do not cover the same failure mode. Enterprises need a way to detect bad outputs after the agent begins interacting with real users, data and tools — especially when nobody reviews the deployment decision first.

                                      An independent agent-evaluation market begins to take shape

                                      The vendor data offers a more encouraging sign: enterprises are adding dedicated evaluation tools, and specialist platforms are gaining ground.

                                      OpenAI’s native evals and traces narrowly led as the primary platform at 18%, followed by Confident AI’s DeepEval at 17% and Braintrust at 15%.

                                      Anthropic’s Claude Console and Workbench held 12%, tied with organizations reporting no dedicated evaluation platform. Three options each held 6%: internally built tools, Promptfoo and LangSmith.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 market landscape

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Braintrust’s primary share increased from 8% in June to 15% in July, the biggest gain by one vendor and the one the report flags as statistically significant. DeepEval rose from 12% to 17%. Use of no purpose-built platform declined five points to 12%, although that smaller change does not by itself confirm a trend.

                                      Because many companies use more than one tool, the broader footprints are larger. OpenAI native evaluation appeared somewhere in 31% of stacks, DeepEval in 27%, Braintrust in 22% and Anthropic’s native tooling in 20%. Custom internal tooling reached 14%, while Weights & Biases Weave and open-source Langfuse each reached 11%.

                                      These are adoption figures, not product-performance scores. The survey does not establish that one vendor produces more reliable agents than another. Still, the results point toward evaluation becoming a distinct enterprise software layer rather than a loose collection of internal scripts or a feature used only inside a model provider’s platform.

                                      Purchasing priorities are changing with that market. The proportion choosing integration ease as the decisive factor climbed 12 points to 39%, displacing cost, which fell from 28% to 23%. Evaluation accuracy ranked second at 28%. The combined average for satisfaction, implementation simplicity and economic value was 3.9 out of five.

                                      The move from price toward integration suggests enterprises increasingly want a tool they can install into existing development and monitoring pipelines now. Yet their leading success metric remains evaluation consistency at 38%, followed by fewer failures and regressions at 20%. Buyers are selecting for fit while still judging results on repeatability.

                                      Switching intent also cooled: 56% still expected to add or replace a platform during the coming year, down from 64% in June.

                                      The proportion staying put increased eight points, reaching 44%. Together with specialist adoption, they suggest some buyers are moving from evaluation to implementation.

                                      Human review is becoming the hedge against automated misses

                                      The budget data reveals how enterprises are managing the contradiction between greater autonomy and unreliable evaluation.

                                      People-centered review workflows edged narrowly ahead of production observability as the most frequently cited area for increased investment, 31% to 30%.

                                      VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026 human review

                                      Credit: VentureBeat Intelligence Agentic Reliability & Evaluations Pulse Tracker July 2026

                                      Automated evaluation pipelines ranked third at 19%, followed by testing for safety and policy compliance at 16%. Only 6% said their reliability and evaluation budget was not increasing.

                                      Among enterprises that had experienced a testing miss, 38% said people-centered review would receive the fastest investment growth, compared with 24% of organizations that had not been burned.

                                      That produces an apparent paradox: the burned group is most likely to remove people from the release checkpoint and most likely to increase spending on people elsewhere in the process. The strategy appears to be automation with a human backstop — allow agents to move faster, then use reviewers to catch what automated evaluation misses.

                                      The open question is whether that model scales. Agent deployments and automated checks can grow with software volume. Reviewer hours do not fall at the same rate. Enterprises may therefore be replacing a human approval step with a larger downstream review function rather than eliminating human oversight.

                                      The narrow but consequential read

                                      July’s data does not show that enterprise agent evaluation is failing everywhere, nor does it prove automated judges are getting worse. It shows something more precise: confidence rose before the measured failure incidence improved.

                                      At the same time, the infrastructure around evaluation is maturing. More enterprises are adopting specialist tools, integration has become the leading purchase criterion and companies that have already experienced failures are increasing investment in human review. The market recognizes the problem and is spending against it.

                                      But the central reliability result remains stubborn. Nearly half of surveyed enterprises still report that an AI feature cleared internal checks before disappointing a customer. Respondents with that experience place less faith in automation — and move faster toward deployments with no human approval.

                                      The report frames this as an incomplete verification model: a passing pre-deployment score marks the start of monitoring, not the end of it. For most enterprises, the production quality checks and evaluation testing that would close that gap still aren’t in place.

                                      ● Canal oficial · Gratis
                                      ¡Recibe las noticias antes que nadie!
                                      Únete a nuestro canal de WhatsApp y mantente informado al instante, sin spam.
                                      Unirme ahora →
                                      ● Noticias al instante ● Cobertura nacional ● Periodismo real Despertar Matinal
                                      — Redacción Despertar Matinal

                                      — Redacción Despertar Matinal

                                      Programa radial que te conecta con la información desde temprano en la mañana.

                                      Next Post
                                      La economía no se entiende mirando por el espejo retrovisor

                                      La economía no se entiende mirando por el espejo retrovisor

                                      Deja una respuesta Cancelar la respuesta

                                      Tu dirección de correo electrónico no será publicada. Los campos obligatorios están marcados con *

                                      Canal de WhatsApp

                                      WhatsApp logo WhatsApp

                                      Canal · Despertar Matinal

                                      Únete a nuestro
                                      Canal

                                      Seguir ahora

                                      El clima

                                      Canal de YouTube

                                      YouTube

                                      Canal · Despertar Matinal

                                      Mira nuestro
                                      Canal

                                      Ver ahora

                                      Escúchanos en Spotify

                                      Spotify

                                      Podcast · Despertar Matinal

                                      Escucha nuestro
                                      Podcast

                                      Escuchar ahora

                                      Noticias Populares

                                      • CRR Las Parras crea talleres industriales de producción de colchones, ropa y tapicería

                                        CRR Las Parras crea talleres industriales de producción de colchones, ropa y tapicería

                                        0 shares
                                        Share 0 Tweet 0
                                      • Milton Morrison reafirma alianza con Abinader y anuncia nueva…

                                        0 shares
                                        Share 0 Tweet 0
                                      • La economía no se entiende mirando por el espejo retrovisor

                                        0 shares
                                        Share 0 Tweet 0
                                      • El CEO de PlayStation habló sobre la fecha de lanzamiento de la PS6

                                        0 shares
                                        Share 0 Tweet 0
                                      • Norway’s King Harald admitted to hospital and put on sick leave

                                        0 shares
                                        Share 0 Tweet 0

                                      Medio digital independiente con análisis, opinión y periodismo responsable desde República Dominicana.

                                      Secciones populares

                                      • Política
                                      • Economía & Negocios
                                      • Justicia
                                      • Turismo
                                      • Tecnología
                                      • Entretenimiento
                                      • Mundo
                                      • Cine y Series
                                      • Música
                                      • Moda

                                      Contenido

                                      • Titulares del Día
                                      • Mundo
                                      • Nacionales
                                      • Política
                                      • Deportes
                                      • Economía & Negocios
                                      • Ciencia
                                      • Entretenimiento
                                      • Podcast
                                      • Opinión
                                      • Despertar Matinal TV
                                      • Editoriales

                                      Corporativo

                                      • Sobre nosotros
                                      • Publicidad
                                      • Sala de prensa
                                      • Contacto
                                      • Política de Privacidad
                                      • Eliminación de Datos

                                      Boletines

                                      Suscríbete a nuestro boletín
                                      Recibe las noticias más importantes cada mañana.

                                      • Nosotros
                                      • Publicidad
                                      • Trabaja con nosotros
                                      • Contactos

                                      © 2025 Despertar Matinal. Aviso Legal - comunícate con nuestra redacción y obtén más información sobre Despertar Matinal..

                                      No Result
                                      View All Result
                                      • Home

                                      © 2025 Despertar Matinal. Aviso Legal - comunícate con nuestra redacción y obtén más información sobre Despertar Matinal..

                                      Welcome Back!

                                      Login to your account below

                                      Forgotten Password?

                                      Retrieve your password

                                      Please enter your username or email address to reset your password.

                                      Log In

                                      Desarrollado por
                                      ►
                                      Las cookies necesarias habilitan funciones esenciales del sitio como inicios de sesión seguros y ajustes de preferencias de consentimiento. No almacenan datos personales.
                                      Ninguno
                                      ►
                                      Las cookies funcionales soportan funciones como compartir contenido en redes sociales, recopilar comentarios y habilitar herramientas de terceros.
                                      Ninguno
                                      ►
                                      Las cookies analíticas rastrean las interacciones de los visitantes, proporcionando información sobre métricas como el número de visitantes, la tasa de rebote y las fuentes de tráfico.
                                      Ninguno
                                      ►
                                      Las cookies de publicidad ofrecen anuncios personalizados basados en tus visitas anteriores y analizan la efectividad de las campañas publicitarias.
                                      Ninguno
                                      ►
                                      Las cookies no clasificadas son aquellas que estamos en proceso de clasificar, junto con los proveedores de cookies individuales.
                                      Ninguno
                                      Desarrollado por