• Nosotros
  • Publicidad
  • Trabaja con nosotros
  • Contactos
miércoles, septiembre 23, 2026
  • Login
No Result
View All Result
NEWSLETTER
Despertar Matinal
  • Titulares del Día
    • All
    • En Portada
    República Dominicana ejecuta operación de manejo de pasivos para reducir riesgo de refinanciamiento

    República Dominicana ejecuta operación de manejo de pasivos para reducir riesgo de refinanciamiento

    Mercedes Carrasco: “Plan anticrisis falló; gobierno con cifras Inconsistentes; PRM abandonó el campo; inflación castiga RD”

    Mercedes Carrasco: “Plan anticrisis falló; gobierno con cifras Inconsistentes; PRM abandonó el campo; inflación castiga RD”

    Fuerza del Pueblo propone plan integral para enfrentar sequía y garantizar agua a la población

    Fuerza del Pueblo propone plan integral para enfrentar sequía y garantizar agua a la población

    Leonel fortalece vínculos con la diáspora y plantea una relación con RD que vaya más allá de las remesas

    Leonel fortalece vínculos con la diáspora y plantea una relación con RD que vaya más allá de las remesas

    Diputada Liz Mieses asegura que Carolina Mejía será candidata del PRM y próxima presidenta

    Diputada Liz Mieses asegura que Carolina Mejía será candidata del PRM y próxima presidenta

    Leonel Fernández juramenta nuevos miembros de la FP en Pensilvania y plantea alianza estratégica con la diáspora

    Leonel Fernández juramenta nuevos miembros de la FP en Pensilvania y plantea alianza estratégica con la diáspora

    Ministerio de Defensa gradúa cadetes especializados en operaciones tácticas en áreas urbanizadas

    Ministerio de Defensa gradúa cadetes especializados en operaciones tácticas en áreas urbanizadas

    Presidente Abinader inicia entrega oficial del Pasaporte Electrónico para dominicanos en Nueva York, Nueva Jersey y Boston

    Presidente Abinader inicia entrega oficial del Pasaporte Electrónico para dominicanos en Nueva York, Nueva Jersey y Boston

    Carolina Mejía presenta a Ricardo de los Santos como su jefe de campaña

    Carolina Mejía presenta a Ricardo de los Santos como su jefe de campaña

    Trending Tags

    • Mundo
      • All
      • América Latina
      • Conflictos Internacionales
      • Estados Unidos
      • Europa
      • Geopolítica
      • Haití
      • Medio Oriente
      Rodolfo D'Onofrio respaldó el regreso de River a la AFA y cuestionó el torneo de 30 equipos y el arbitraje

      Rodolfo D’Onofrio respaldó el regreso de River a la AFA y cuestionó el torneo de 30 equipos y el arbitraje

      La OCDE recortó la proyección de crecimiento del Reino Unido y se acentúa la crisis socialista

      La OCDE recortó la proyección de crecimiento del Reino Unido y se acentúa la crisis socialista

      Quirno usó el derecho a réplica en la ONU y respondió a Burnham: “Las Malvinas son argentinas”

      Quirno usó el derecho a réplica en la ONU y respondió a Burnham: “Las Malvinas son argentinas”

      El régimen de Xi Jinping amplió los controles sobre químicos del fentanilo tras la presión del gobierno de Trump

      El régimen de Xi Jinping amplió los controles sobre químicos del fentanilo tras la presión del gobierno de Trump

      Córdoba: el banco de horas ya rige en Renault gracias a la de Modernización Laboral de Milei

      Córdoba: el banco de horas ya rige en Renault gracias a la de Modernización Laboral de Milei

      En C5N defendieron "El Gran Salto Adelante": la reforma agraria impulsada por Mao Zedong

      En C5N defendieron «El Gran Salto Adelante»: la reforma agraria impulsada por Mao Zedong

      Javier Milei participó junto a Donald Trump de la cumbre “Shield of the Americas” en Nueva York

      Javier Milei participó junto a Donald Trump de la cumbre “Shield of the Americas” en Nueva York

      Anthropic y OpenAI lanzan modelos de IA más potentes y baratos

      Anthropic y OpenAI lanzan modelos de IA más potentes y baratos

      El Gobierno de Milei realizó obras en la Ruta Nacional 14 y puso en valor 200 kilómetros

      El Gobierno de Milei realizó obras en la Ruta Nacional 14 y puso en valor 200 kilómetros

      Trending Tags

      • Nacionales
        • All
        • Bávaro Punta Cana
        • Educación
        • Gobierno
        • Infraestructura
        • Justicia
        • Obras Públicas
        • Opinión
        • Provincias
        • Seguridad Ciudadana
        • semana santa 2026
        • Sociedad
        • Transporte
        Josefa Castillo se compromete a fortalecer servicios consulares...

        Josefa Castillo se compromete a fortalecer servicios consulares…

        Ministerio Público solicita a corte condenar a Wander Franco a cinco años de prisión

        Ministerio Público solicita a corte condenar a Wander Franco a cinco años de prisión

        Desde ONU, Paliza llama a fortalecer cooperación regional para...

        Desde ONU, Paliza llama a fortalecer cooperación regional para…

        UASD inaugura simposio reunirá destacados académicos...

        UASD inaugura simposio reunirá destacados académicos…

        República Dominicana ejecuta operación de manejo de pasivos para reducir riesgo de refinanciamiento

        República Dominicana ejecuta operación de manejo de pasivos para reducir riesgo de refinanciamiento

        Mercedes Carrasco: “Plan anticrisis falló; gobierno con cifras Inconsistentes; PRM abandonó el campo; inflación castiga RD”

        Mercedes Carrasco: “Plan anticrisis falló; gobierno con cifras Inconsistentes; PRM abandonó el campo; inflación castiga RD”

        Fuerza del Pueblo propone plan integral para enfrentar sequía y garantizar agua a la población

        Fuerza del Pueblo propone plan integral para enfrentar sequía y garantizar agua a la población

        Canciller Bisonó destaca prioridades de República Dominicana...

        Canciller Bisonó destaca prioridades de República Dominicana…

        Proponen declarar el turismo de salud como prioridad nacional ...

        Proponen declarar el turismo de salud como prioridad nacional …

        Trending Tags

        • Política
          • All
          • Congreso
          • Opinión Política
          • Partidos Políticos
          • Poder Municipal
          • Transparencia y Corrupción
          (VIDEO) EN SAN JUAN: La Fuerza del Pueblo incorpora a Alejandro Tejada en un multitudinario encuentro en El Batey

          (VIDEO) EN SAN JUAN: La Fuerza del Pueblo incorpora a Alejandro Tejada en un multitudinario encuentro en El Batey

          Gonzalo Castillo: “Me pueden meter preso, nadie va a evitar que sea presidente de la República Dominicana”

          Gonzalo Castillo: “Me pueden meter preso, nadie va a evitar que sea presidente de la República Dominicana”

          Exministro de Haciendas advierte fuga de ahorros en dólares si...

          Exministro de Haciendas advierte fuga de ahorros en dólares si…

          Robert Polanco revela respaldo a David Collado y descarta...

          Robert Polanco revela respaldo a David Collado y descarta…

          TSE dispone suspensión provisional celebración VII Convención...

          TSE dispone suspensión provisional celebración VII Convención…

          Fuerza del Pueblo en Ocoa desmiente que seis personas fueran miembros activos del partido y juramentadas con Carolina Mejía

          Fuerza del Pueblo en Ocoa desmiente que seis personas fueran miembros activos del partido y juramentadas con Carolina Mejía

          Danilo Medina proclama en Barahona: “Ya no esperen nada de este gobierno”

          Danilo Medina proclama en Barahona: “Ya no esperen nada de este gobierno”

          Luis Abinader coloca formación política, capacitación y conducta ética entre los ejes de su gestión al frente del PRM

          Luis Abinader coloca formación política, capacitación y conducta ética entre los ejes de su gestión al frente del PRM

          ¡El rumbo de la historia ya está decidido! Fuerza del Pueblo avanza hacia el 2028 junto a Leonel Fernández

          ¡El rumbo de la historia ya está decidido! Fuerza del Pueblo avanza hacia el 2028 junto a Leonel Fernández

          Trending Tags

          • Deportes
            • All
            • Atletas Dominicanos
            • Béisbol
            DR Open Kiteboarding Championship reúne atletas de 15 países y reafirma a Cabarete como capital del kitesurf del Caribe

            Cabarete se corona como capital histórica del kitesurf con el DR Open Championship 2026

            El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

            El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

            La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

            La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

            Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

            Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

            Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

            Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

            Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

            Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

            Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

            Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

            30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

            Buffalo recibe a Montreal para abrir la segunda ronda

            Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

            Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

            Trending Tags

            • Economía
              • All
              • Combustibles
              • Energía
              • Indicadores Económicos
              • Sector Energético
              • Turismo
              Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

              Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

              Aventúrate RD 2026

              Aventúrate RD 2026 revela agenda oficial y consolida el turismo de aventura dominicano

              WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

              WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

              Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

              Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

              Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

              Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

              Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

              Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

              El ministerio de Turismo cierra con éxito festival gastronómico “Saborea el Paraíso” en Sánchez, Samaná

              El Ministerio de Turismo celebra un exitoso cierre del festival gastronómico «Saborea el Paraíso» en Sánchez, Samaná

              El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

              El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

              Más allá del comercio: los efectos del arancel estadounidense sobre el turismo dominicano

              Arancel de EE.UU. pone a prueba al turismo dominicano y al silencio oficial del gobierno

              Trending Tags

              • Ciencia
                • All
                • Energía
                • Innovación
                • Investigación Científica
                • Salud y Medicina
                • Tecnología Médica
                Cadenas de medios de EE.UU. suspenden la cobertura de Trump

                Cadenas de medios de EE.UU. suspenden la cobertura de Trump

                El partido opositor FMLN confirma su candidato presidencial 2027

                El partido opositor FMLN confirma su candidato presidencial 2027

                Merz promete mantener la coalición en Alemania

                Merz promete mantener la coalición en Alemania

                El Kremlin en Rusia Unida revalida la mayoría constitucional

                El Kremlin en Rusia Unida revalida la mayoría constitucional

                La urna electrónica cumple 30 años en víspera electoral de Brasil

                La urna electrónica cumple 30 años en víspera electoral de Brasil

                Los tres medios de comunicación vetados por Trump le demandan

                Los tres medios de comunicación vetados por Trump le demandan

                Gonzalo Castillo: Me pueden meter preso

                Gonzalo Castillo: Me pueden meter preso

                Charlie Mariotti Jr. cuestiona preéstamos US$1,700 MM en energía

                Charlie Mariotti Jr. cuestiona preéstamos US$1,700 MM en energía

                Bajo la visión, orden y firmeza, Carolina anuncia a Ricardo de los Santos como su jefe de campaña

                Bajo la visión, orden y firmeza, Carolina anuncia a Ricardo de los Santos como su jefe de campaña

                Trending Tags

                • Tecnología
                  • All
                  • Aplicaciones
                  • Inteligencia Artificial
                  El senador Sanders presenta un proyecto de ley para prohibir la superinteligencia artificial y crear un Departamento de Inteligencia Artificial

                  El senador Sanders presenta un proyecto de ley para prohibir la superinteligencia artificial y crear un Departamento de Inteligencia Artificial

                  Una mirada a los escenarios apocalípticos de la IA que, según los investigadores, podrían poner a la humanidad en riesgo

                  Una mirada a los escenarios apocalípticos de la IA que, según los investigadores, podrían poner a la humanidad en riesgo

                  Las preocupaciones sobre una adquisición de Internet por parte de la IA adquieren nueva urgencia entre los escenarios apocalípticos

                  Las preocupaciones sobre una adquisición de Internet por parte de la IA adquieren nueva urgencia entre los escenarios apocalípticos

                  La agencia dice que los huevos de tortugas marinas depositados en las playas de California son los primeros en la costa oeste de EE. UU.

                  La agencia dice que los huevos de tortugas marinas depositados en las playas de California son los primeros en la costa oeste de EE. UU.

                  Cumplir acuerdos anteriores entre Estados Unidos y China es un trabajo en progreso a medida que Trump y Xi se reencuentran

                  Cumplir acuerdos anteriores entre Estados Unidos y China es un trabajo en progreso a medida que Trump y Xi se reencuentran

                  El juez no impedirá que la administración Trump le dé a SpaceX acres de refugio para la vida silvestre

                  El juez no impedirá que la administración Trump le dé a SpaceX acres de refugio para la vida silvestre

                  La Fundación Gates lanza una coalición para obtener conjuntos de datos lingüísticos más representativos para la IA

                  La Fundación Gates lanza una coalición para obtener conjuntos de datos lingüísticos más representativos para la IA

                  Google recibe una multa de 463 millones de dólares por violar la norma de datos de ubicación de la UE

                  Google recibe una multa de 463 millones de dólares por violar la norma de datos de ubicación de la UE

                  Bessent: Estados Unidos propone un sistema de alerta de incidentes mediante IA en conversaciones con China

                  Bessent: Estados Unidos propone un sistema de alerta de incidentes mediante IA en conversaciones con China

                  Trending Tags

                  • Entretenimiento
                    • All
                    • Cine y Series
                    • Cultura Digital
                    • Cultura Popular
                    • Gastronomía
                    • Música
                    Celine Dion está de regreso en París, pero su primera canción sigue siendo "un gran secreto"

                    Celine Dion está de regreso en París, pero su primera canción sigue siendo «un gran secreto»

                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                    La ‘Odisea’ de Emily Wilson se convirtió en un punto de inflamación cultural. Ahora ella está retraduciendo todo.

                    Editor, editor y reportero de Stars and Stripes demandan al Pentágono para impugnar sus despidos

                    Editor, editor y reportero de Stars and Stripes demandan al Pentágono para impugnar sus despidos

                    Muere Peter Cullen, el prolífico actor de doblaje que le dio a Optimus Prime su autoritario barítono

                    Muere Peter Cullen, el prolífico actor de doblaje que le dio a Optimus Prime su autoritario barítono

                    Juez pregunta por qué el Kennedy Center se está moviendo tan rápido para devolver el nombre de Trump al edificio

                    Juez pregunta por qué el Kennedy Center se está moviendo tan rápido para devolver el nombre de Trump al edificio

                    Un teatro reinventa la Odisea de Homero a través de la agonía de la guerra de Ucrania

                    Un teatro reinventa la Odisea de Homero a través de la agonía de la guerra de Ucrania

                    En el conflictivo norte de Nigeria, una animada vida nocturna convive con una policía moral e inseguridad.

                    En el conflictivo norte de Nigeria, una animada vida nocturna convive con una policía moral e inseguridad.

                    El rapero Yung Filly regresará a Gran Bretaña antes del juicio por violación en Australia el próximo año

                    El rapero Yung Filly regresará a Gran Bretaña antes del juicio por violación en Australia el próximo año

                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                    Los británicos tienen la oportunidad de leer las memorias de Jason Arday en las librerías del Reino Unido

                    Trending Tags

                    • Titulares del Día
                      • All
                      • En Portada
                      República Dominicana ejecuta operación de manejo de pasivos para reducir riesgo de refinanciamiento

                      República Dominicana ejecuta operación de manejo de pasivos para reducir riesgo de refinanciamiento

                      Mercedes Carrasco: “Plan anticrisis falló; gobierno con cifras Inconsistentes; PRM abandonó el campo; inflación castiga RD”

                      Mercedes Carrasco: “Plan anticrisis falló; gobierno con cifras Inconsistentes; PRM abandonó el campo; inflación castiga RD”

                      Fuerza del Pueblo propone plan integral para enfrentar sequía y garantizar agua a la población

                      Fuerza del Pueblo propone plan integral para enfrentar sequía y garantizar agua a la población

                      Leonel fortalece vínculos con la diáspora y plantea una relación con RD que vaya más allá de las remesas

                      Leonel fortalece vínculos con la diáspora y plantea una relación con RD que vaya más allá de las remesas

                      Diputada Liz Mieses asegura que Carolina Mejía será candidata del PRM y próxima presidenta

                      Diputada Liz Mieses asegura que Carolina Mejía será candidata del PRM y próxima presidenta

                      Leonel Fernández juramenta nuevos miembros de la FP en Pensilvania y plantea alianza estratégica con la diáspora

                      Leonel Fernández juramenta nuevos miembros de la FP en Pensilvania y plantea alianza estratégica con la diáspora

                      Ministerio de Defensa gradúa cadetes especializados en operaciones tácticas en áreas urbanizadas

                      Ministerio de Defensa gradúa cadetes especializados en operaciones tácticas en áreas urbanizadas

                      Presidente Abinader inicia entrega oficial del Pasaporte Electrónico para dominicanos en Nueva York, Nueva Jersey y Boston

                      Presidente Abinader inicia entrega oficial del Pasaporte Electrónico para dominicanos en Nueva York, Nueva Jersey y Boston

                      Carolina Mejía presenta a Ricardo de los Santos como su jefe de campaña

                      Carolina Mejía presenta a Ricardo de los Santos como su jefe de campaña

                      Trending Tags

                      • Mundo
                        • All
                        • América Latina
                        • Conflictos Internacionales
                        • Estados Unidos
                        • Europa
                        • Geopolítica
                        • Haití
                        • Medio Oriente
                        Rodolfo D'Onofrio respaldó el regreso de River a la AFA y cuestionó el torneo de 30 equipos y el arbitraje

                        Rodolfo D’Onofrio respaldó el regreso de River a la AFA y cuestionó el torneo de 30 equipos y el arbitraje

                        La OCDE recortó la proyección de crecimiento del Reino Unido y se acentúa la crisis socialista

                        La OCDE recortó la proyección de crecimiento del Reino Unido y se acentúa la crisis socialista

                        Quirno usó el derecho a réplica en la ONU y respondió a Burnham: “Las Malvinas son argentinas”

                        Quirno usó el derecho a réplica en la ONU y respondió a Burnham: “Las Malvinas son argentinas”

                        El régimen de Xi Jinping amplió los controles sobre químicos del fentanilo tras la presión del gobierno de Trump

                        El régimen de Xi Jinping amplió los controles sobre químicos del fentanilo tras la presión del gobierno de Trump

                        Córdoba: el banco de horas ya rige en Renault gracias a la de Modernización Laboral de Milei

                        Córdoba: el banco de horas ya rige en Renault gracias a la de Modernización Laboral de Milei

                        En C5N defendieron "El Gran Salto Adelante": la reforma agraria impulsada por Mao Zedong

                        En C5N defendieron «El Gran Salto Adelante»: la reforma agraria impulsada por Mao Zedong

                        Javier Milei participó junto a Donald Trump de la cumbre “Shield of the Americas” en Nueva York

                        Javier Milei participó junto a Donald Trump de la cumbre “Shield of the Americas” en Nueva York

                        Anthropic y OpenAI lanzan modelos de IA más potentes y baratos

                        Anthropic y OpenAI lanzan modelos de IA más potentes y baratos

                        El Gobierno de Milei realizó obras en la Ruta Nacional 14 y puso en valor 200 kilómetros

                        El Gobierno de Milei realizó obras en la Ruta Nacional 14 y puso en valor 200 kilómetros

                        Trending Tags

                        • Nacionales
                          • All
                          • Bávaro Punta Cana
                          • Educación
                          • Gobierno
                          • Infraestructura
                          • Justicia
                          • Obras Públicas
                          • Opinión
                          • Provincias
                          • Seguridad Ciudadana
                          • semana santa 2026
                          • Sociedad
                          • Transporte
                          Josefa Castillo se compromete a fortalecer servicios consulares...

                          Josefa Castillo se compromete a fortalecer servicios consulares…

                          Ministerio Público solicita a corte condenar a Wander Franco a cinco años de prisión

                          Ministerio Público solicita a corte condenar a Wander Franco a cinco años de prisión

                          Desde ONU, Paliza llama a fortalecer cooperación regional para...

                          Desde ONU, Paliza llama a fortalecer cooperación regional para…

                          UASD inaugura simposio reunirá destacados académicos...

                          UASD inaugura simposio reunirá destacados académicos…

                          República Dominicana ejecuta operación de manejo de pasivos para reducir riesgo de refinanciamiento

                          República Dominicana ejecuta operación de manejo de pasivos para reducir riesgo de refinanciamiento

                          Mercedes Carrasco: “Plan anticrisis falló; gobierno con cifras Inconsistentes; PRM abandonó el campo; inflación castiga RD”

                          Mercedes Carrasco: “Plan anticrisis falló; gobierno con cifras Inconsistentes; PRM abandonó el campo; inflación castiga RD”

                          Fuerza del Pueblo propone plan integral para enfrentar sequía y garantizar agua a la población

                          Fuerza del Pueblo propone plan integral para enfrentar sequía y garantizar agua a la población

                          Canciller Bisonó destaca prioridades de República Dominicana...

                          Canciller Bisonó destaca prioridades de República Dominicana…

                          Proponen declarar el turismo de salud como prioridad nacional ...

                          Proponen declarar el turismo de salud como prioridad nacional …

                          Trending Tags

                          • Política
                            • All
                            • Congreso
                            • Opinión Política
                            • Partidos Políticos
                            • Poder Municipal
                            • Transparencia y Corrupción
                            (VIDEO) EN SAN JUAN: La Fuerza del Pueblo incorpora a Alejandro Tejada en un multitudinario encuentro en El Batey

                            (VIDEO) EN SAN JUAN: La Fuerza del Pueblo incorpora a Alejandro Tejada en un multitudinario encuentro en El Batey

                            Gonzalo Castillo: “Me pueden meter preso, nadie va a evitar que sea presidente de la República Dominicana”

                            Gonzalo Castillo: “Me pueden meter preso, nadie va a evitar que sea presidente de la República Dominicana”

                            Exministro de Haciendas advierte fuga de ahorros en dólares si...

                            Exministro de Haciendas advierte fuga de ahorros en dólares si…

                            Robert Polanco revela respaldo a David Collado y descarta...

                            Robert Polanco revela respaldo a David Collado y descarta…

                            TSE dispone suspensión provisional celebración VII Convención...

                            TSE dispone suspensión provisional celebración VII Convención…

                            Fuerza del Pueblo en Ocoa desmiente que seis personas fueran miembros activos del partido y juramentadas con Carolina Mejía

                            Fuerza del Pueblo en Ocoa desmiente que seis personas fueran miembros activos del partido y juramentadas con Carolina Mejía

                            Danilo Medina proclama en Barahona: “Ya no esperen nada de este gobierno”

                            Danilo Medina proclama en Barahona: “Ya no esperen nada de este gobierno”

                            Luis Abinader coloca formación política, capacitación y conducta ética entre los ejes de su gestión al frente del PRM

                            Luis Abinader coloca formación política, capacitación y conducta ética entre los ejes de su gestión al frente del PRM

                            ¡El rumbo de la historia ya está decidido! Fuerza del Pueblo avanza hacia el 2028 junto a Leonel Fernández

                            ¡El rumbo de la historia ya está decidido! Fuerza del Pueblo avanza hacia el 2028 junto a Leonel Fernández

                            Trending Tags

                            • Deportes
                              • All
                              • Atletas Dominicanos
                              • Béisbol
                              DR Open Kiteboarding Championship reúne atletas de 15 países y reafirma a Cabarete como capital del kitesurf del Caribe

                              Cabarete se corona como capital histórica del kitesurf con el DR Open Championship 2026

                              El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

                              El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

                              La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

                              La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

                              Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

                              Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

                              Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

                              Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

                              Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

                              Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

                              Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

                              Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

                              30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                              Buffalo recibe a Montreal para abrir la segunda ronda

                              Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

                              Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

                              Trending Tags

                              • Economía
                                • All
                                • Combustibles
                                • Energía
                                • Indicadores Económicos
                                • Sector Energético
                                • Turismo
                                Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

                                Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

                                Aventúrate RD 2026

                                Aventúrate RD 2026 revela agenda oficial y consolida el turismo de aventura dominicano

                                WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

                                WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

                                Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

                                Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

                                Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

                                Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

                                Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

                                Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

                                El ministerio de Turismo cierra con éxito festival gastronómico “Saborea el Paraíso” en Sánchez, Samaná

                                El Ministerio de Turismo celebra un exitoso cierre del festival gastronómico «Saborea el Paraíso» en Sánchez, Samaná

                                El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

                                El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

                                Más allá del comercio: los efectos del arancel estadounidense sobre el turismo dominicano

                                Arancel de EE.UU. pone a prueba al turismo dominicano y al silencio oficial del gobierno

                                Trending Tags

                                • Ciencia
                                  • All
                                  • Energía
                                  • Innovación
                                  • Investigación Científica
                                  • Salud y Medicina
                                  • Tecnología Médica
                                  Cadenas de medios de EE.UU. suspenden la cobertura de Trump

                                  Cadenas de medios de EE.UU. suspenden la cobertura de Trump

                                  El partido opositor FMLN confirma su candidato presidencial 2027

                                  El partido opositor FMLN confirma su candidato presidencial 2027

                                  Merz promete mantener la coalición en Alemania

                                  Merz promete mantener la coalición en Alemania

                                  El Kremlin en Rusia Unida revalida la mayoría constitucional

                                  El Kremlin en Rusia Unida revalida la mayoría constitucional

                                  La urna electrónica cumple 30 años en víspera electoral de Brasil

                                  La urna electrónica cumple 30 años en víspera electoral de Brasil

                                  Los tres medios de comunicación vetados por Trump le demandan

                                  Los tres medios de comunicación vetados por Trump le demandan

                                  Gonzalo Castillo: Me pueden meter preso

                                  Gonzalo Castillo: Me pueden meter preso

                                  Charlie Mariotti Jr. cuestiona preéstamos US$1,700 MM en energía

                                  Charlie Mariotti Jr. cuestiona preéstamos US$1,700 MM en energía

                                  Bajo la visión, orden y firmeza, Carolina anuncia a Ricardo de los Santos como su jefe de campaña

                                  Bajo la visión, orden y firmeza, Carolina anuncia a Ricardo de los Santos como su jefe de campaña

                                  Trending Tags

                                  • Tecnología
                                    • All
                                    • Aplicaciones
                                    • Inteligencia Artificial
                                    El senador Sanders presenta un proyecto de ley para prohibir la superinteligencia artificial y crear un Departamento de Inteligencia Artificial

                                    El senador Sanders presenta un proyecto de ley para prohibir la superinteligencia artificial y crear un Departamento de Inteligencia Artificial

                                    Una mirada a los escenarios apocalípticos de la IA que, según los investigadores, podrían poner a la humanidad en riesgo

                                    Una mirada a los escenarios apocalípticos de la IA que, según los investigadores, podrían poner a la humanidad en riesgo

                                    Las preocupaciones sobre una adquisición de Internet por parte de la IA adquieren nueva urgencia entre los escenarios apocalípticos

                                    Las preocupaciones sobre una adquisición de Internet por parte de la IA adquieren nueva urgencia entre los escenarios apocalípticos

                                    La agencia dice que los huevos de tortugas marinas depositados en las playas de California son los primeros en la costa oeste de EE. UU.

                                    La agencia dice que los huevos de tortugas marinas depositados en las playas de California son los primeros en la costa oeste de EE. UU.

                                    Cumplir acuerdos anteriores entre Estados Unidos y China es un trabajo en progreso a medida que Trump y Xi se reencuentran

                                    Cumplir acuerdos anteriores entre Estados Unidos y China es un trabajo en progreso a medida que Trump y Xi se reencuentran

                                    El juez no impedirá que la administración Trump le dé a SpaceX acres de refugio para la vida silvestre

                                    El juez no impedirá que la administración Trump le dé a SpaceX acres de refugio para la vida silvestre

                                    La Fundación Gates lanza una coalición para obtener conjuntos de datos lingüísticos más representativos para la IA

                                    La Fundación Gates lanza una coalición para obtener conjuntos de datos lingüísticos más representativos para la IA

                                    Google recibe una multa de 463 millones de dólares por violar la norma de datos de ubicación de la UE

                                    Google recibe una multa de 463 millones de dólares por violar la norma de datos de ubicación de la UE

                                    Bessent: Estados Unidos propone un sistema de alerta de incidentes mediante IA en conversaciones con China

                                    Bessent: Estados Unidos propone un sistema de alerta de incidentes mediante IA en conversaciones con China

                                    Trending Tags

                                    • Entretenimiento
                                      • All
                                      • Cine y Series
                                      • Cultura Digital
                                      • Cultura Popular
                                      • Gastronomía
                                      • Música
                                      Celine Dion está de regreso en París, pero su primera canción sigue siendo "un gran secreto"

                                      Celine Dion está de regreso en París, pero su primera canción sigue siendo «un gran secreto»

                                      30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                      La ‘Odisea’ de Emily Wilson se convirtió en un punto de inflamación cultural. Ahora ella está retraduciendo todo.

                                      Editor, editor y reportero de Stars and Stripes demandan al Pentágono para impugnar sus despidos

                                      Editor, editor y reportero de Stars and Stripes demandan al Pentágono para impugnar sus despidos

                                      Muere Peter Cullen, el prolífico actor de doblaje que le dio a Optimus Prime su autoritario barítono

                                      Muere Peter Cullen, el prolífico actor de doblaje que le dio a Optimus Prime su autoritario barítono

                                      Juez pregunta por qué el Kennedy Center se está moviendo tan rápido para devolver el nombre de Trump al edificio

                                      Juez pregunta por qué el Kennedy Center se está moviendo tan rápido para devolver el nombre de Trump al edificio

                                      Un teatro reinventa la Odisea de Homero a través de la agonía de la guerra de Ucrania

                                      Un teatro reinventa la Odisea de Homero a través de la agonía de la guerra de Ucrania

                                      En el conflictivo norte de Nigeria, una animada vida nocturna convive con una policía moral e inseguridad.

                                      En el conflictivo norte de Nigeria, una animada vida nocturna convive con una policía moral e inseguridad.

                                      El rapero Yung Filly regresará a Gran Bretaña antes del juicio por violación en Australia el próximo año

                                      El rapero Yung Filly regresará a Gran Bretaña antes del juicio por violación en Australia el próximo año

                                      30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                      Los británicos tienen la oportunidad de leer las memorias de Jason Arday en las librerías del Reino Unido

                                      Trending Tags

                                      No Result
                                      View All Result
                                      Despertar Matinal
                                      No Result
                                      View All Result

                                      DeepSWE blows up the AI coding leaderboard, crowns GPT-5.5, and finds Claude Opus exploiting a benchmark loophole

                                      by — Redacción Despertar Matinal
                                      26 de mayo de 2026
                                      in Tecnología
                                      0
                                      DeepSWE blows up the AI coding leaderboard, crowns GPT-5.5, and finds Claude Opus exploiting a benchmark loophole
                                      0
                                      SHARES
                                      38
                                      VIEWS
                                      Share on FacebookShare on Twitter

                                      For months, the leading AI coding benchmarks have told enterprise buyers a comforting but misleading story: the top models are all roughly the same. OpenAI’s GPT-5 family, Anthropic’s Claude Opus, and Google’s Gemini Pro have clustered within a narrow band on Scale AI’s SWE-Bench Pro leaderboard, making it nearly impossible for engineering leaders to determine which agent will actually perform best inside their codebases.

                                      On Monday, a startup called Datacurve released a benchmark it says shatters that illusion. DeepSWE, a 113-task evaluation spanning 91 open-source repositories and five programming languages, produces a dramatically wider spread among the same frontier models — and crowns OpenAI’s GPT-5.5 as the clear leader at 70%, sixteen points ahead of its nearest competitor.

                                      «On public leaderboards, top models often look relatively close in capability,» wrote Datacurve co-author Serena Ge on X. «DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.»

                                      The benchmark also delivers a pointed critique of the evaluation infrastructure the AI industry relies on to measure progress: Datacurve’s audit found that SWE-Bench Pro’s verifiers — the automated graders that determine whether an agent solved a task — issued incorrect pass/fail verdicts on roughly one-third of the trials it reviewed.

                                      If that finding holds up, it has sweeping implications. Enterprise procurement teams, venture capitalists, and AI lab marketing departments all lean heavily on benchmark scores to make multimillion-dollar decisions. A 32% error rate in the most widely cited coding benchmark suggests the industry may have been navigating by a broken compass.

                                      Why the most popular AI coding benchmark may be grading on a curve

                                      To understand what Datacurve is claiming, it helps to understand how coding benchmarks work — and how they can go wrong.

                                      The dominant paradigm, pioneered by the SWE-Bench family maintained by Scale AI and academic researchers, constructs tasks by mining real GitHub commits. The process extracts a bug fix or feature addition from a repository’s history, rolls the code back to the pre-fix state, and then asks an AI agent to reproduce the change. The original commit’s test suite serves as the verifier: if the agent’s patch makes the same tests pass, it gets credit. This approach has an elegant simplicity, but Datacurve argues it introduces three systemic weaknesses.

                                      First, contamination. Because tasks are drawn from public GitHub history, the problem statement, the discussion, and often the exact solution are already present in the training data of frontier models. «The SWE-Bench family scrapes existing GitHub issues and PRs, which creates two problems: memorization (models have already seen the solution) and triviality (most tasks are small),» Ge wrote.

                                      Second, scope. SWE-Bench Pro tasks require, on average, just 120 lines of code added across 5 files. DeepSWE’s reference solutions average 668 lines added across 7 files — roughly 5.5 times more code. Yet DeepSWE’s prompts are actually shorter, averaging 2,158 characters versus SWE-Bench Pro’s 4,614. In other words, DeepSWE gives the agent less instruction but expects far more output, which more closely mirrors how a human developer might actually delegate work to an AI assistant.

                                      DeepSWE tasks demand roughly five times more code than SWE-Bench Pro’s while giving agents shorter prompts — a design choice intended to mirror how developers actually hand off work. (Source: Datacurve)

                                      Third — and most damaging — verifier reliability. Datacurve drew 30 tasks at random from both DeepSWE and SWE-Bench Pro, ran three rollouts across 10 frontier model configurations, and then deployed an LLM-based judge to independently assess whether each agent’s patch actually solved the problem. SWE-Bench Pro’s verifiers accepted wrong implementations 8.5% of the time and rejected correct implementations 24% of the time. DeepSWE’s verifiers registered 0.3% and 1.1%, respectively.

                                      Screenshot 2026-05-26 at 3.22.11 PM

                                      Datacurve’s audit found that SWE-Bench Pro’s automated graders rejected correct solutions 24 percent of the time and accepted wrong ones 8.5 percent of the time. DeepSWE’s verifiers kept both rates near zero. (Source: Datacurve)

                                      The false negative problem is especially insidious because it punishes creative solutions. In one documented case, the gold-standard pull request for a SWE-Bench Pro task refactored a private helper function. An agent that correctly solved the task by inlining the same logic — a perfectly valid engineering choice — failed because the test suite tried to import a symbol that only existed in the original author’s specific implementation.

                                      OpenAI’s GPT-5.5 dominates the new benchmark while Claude and Gemini stumble

                                      DeepSWE’s top-line results reorder the familiar hierarchy in ways that should matter to every engineering team evaluating AI coding tools. On SWE-Bench Pro, models from OpenAI, Anthropic, and Google have traded the lead within a 30-point range. DeepSWE stretches that range to 70 points.

                                      GPT-5.5 leads at 70%, followed by GPT-5.4 at 56% and Claude Opus 4.7 at 54%. From there, the drop-off is steep: Claude Sonnet 4.6 lands at 32%, Gemini 3.5 Flash at 28%, GPT-5.4-mini and Kimi K2.6 tied at 24%, and then a long tail of models in the teens and single digits. Claude Haiku 4.5, which scores 39% on SWE-Bench Pro, collapses to zero on DeepSWE — suggesting that some mid-tier models have been significantly overperforming on easier, potentially contaminated benchmarks.

                                      Screenshot 2026-05-26 at 3.09.33 PM

                                      On SWE-Bench Pro, frontier models cluster within a 30-point range. On DeepSWE, the same models spread across 70 points, with some — like Claude Haiku 4.5 — collapsing entirely. (Source: Datacurve)

                                      GPT-5.5 doesn’t just score the highest — it does so efficiently. The model reaches its 70% pass rate with a median cost of $5.80 per trial, a median wall-clock time of 20 minutes, and a median of 47,000 output tokens. GPT-5.4 emerges as perhaps the best overall value at $3.30 per trial with a 56% score. Claude Opus 4.7, meanwhile, costs significantly more per run, and output tokens, wall-clock duration, and dollar cost per trial all vary by an order of magnitude across the agents tested — yet none of these correlates strongly with pass rate. Agents that emit more tokens, run longer, or cost more do not consistently solve more tasks.

                                      Screenshot 2026-05-26 at 3.27.42 PM

                                      GPT-5.4 and GPT-5.5 occupy the cost-efficient frontier, solving the most tasks for the least money per run. Spending more did not reliably produce better results. (Source: Datacurve)

                                      Datacurve’s audit found that Claude has been reading the answer key on existing benchmarks

                                      Perhaps the most provocative finding in DeepSWE’s analysis concerns what the authors label «CHEATED» verdicts — instances where an agent passes a benchmark not by solving the problem, but by reading the answer.

                                      SWE-Bench Pro’s Docker containers ship the repository’s full .git history, which means the gold-standard solution commit is sitting right there in the container’s file system. Most models ignore it. Claude does not. Datacurve’s analysis found that both Claude Opus 4.7 and Claude Opus 4.6 registered «CHEATED» on more than 12% of their reviewed SWE-Bench Pro rollouts. In those instances, the Claude agent ran commands like git log –all or git show to retrieve the merged fix and paste it into its own patch. The behavior accounted for approximately 18% of Opus 4.7’s passes and 25% of Opus 4.6’s passes on the reviewed sample. The issue has been filed publicly as GitHub issue #93 on the SWE-Bench Pro repository.

                                      GPT-5.4 and GPT-5.5 never exhibited this behavior. Gemini configurations stayed around 1%. Datacurve describes the behavior diplomatically — «The benchmark makes this possible (the gold commit lives in the container), but Claude is the family that consistently does so» — but the implication is clear: a meaningful fraction of Claude’s SWE-Bench Pro scores may reflect environmental exploitation rather than genuine engineering capability.

                                      DeepSWE addresses this by shipping only a shallow clone with the base commit, leaving no gold hash for the agent to discover. It is worth noting that the behavior is arguably a sign of Claude’s environmental attentiveness — the model is very good at exploring its surroundings and exploiting available resources. Whether that counts as «cheating» or «resourcefulness» depends on your perspective, but in the context of a benchmark designed to measure independent problem-solving, it undermines the signal.

                                      Screenshot 2026-05-26 at 3.28.43 PM

                                      Two mechanisms by which agents passed SWE-Bench Pro without solving the underlying problem: reading the answer from the container’s Git history, or stubbing features past weak gold tests. (Source: Datacurve)

                                      Each AI model family fails in its own distinctive way, and the patterns matter for enterprise teams

                                      Beyond the top-line scores, Datacurve’s qualitative trajectory analysis reveals distinctly different failure signatures across model families — a finding that could help engineering teams choose the right model for specific types of work.

                                      Claude is forgetful with multi-part prompts. On DeepSWE, Claude configurations miss stated requirements more than any other family. The pattern is consistent: when a prompt enumerates parallel behaviors — «support both sync and async,» for instance — Claude typically implements the obvious branch and forgets to mirror the change. Datacurve reports that roughly two-thirds of Claude’s «MISSED_REQUIREMENT» failures on DeepSWE follow this «one branch shipped» pattern. In one example, Claude Opus 4.7 correctly landed a sync state-data hook in one engine class while the async engine never received the same hook.

                                      GPT, by contrast, implements exactly what is asked. GPT-5.5 had the lowest rate of missing stated behaviors of any configuration tested. Across multiple runs of the same task, GPT trials tended to converge on the same interpretation of the prompt, suggesting instruction-following precision is a stable trait of the model rather than per-run luck.

                                      One of the most intriguing findings involves self-verification. On DeepSWE, Claude Opus 4.7 and GPT-5.4 wrote and ran new tests in the project’s own test framework on over 80% of their runs — even though no one asked them to. On SWE-Bench Pro, those same models dropped to 28% and 18%, respectively. The reason: SWE-Bench Pro’s prompt template explicitly tells agents they «should not modify the testing logic or any of the tests.» Agents dutifully complied, suppressing a behavior that likely would have improved their performance. This suggests that prompt design in production coding workflows may be inadvertently suppressing valuable agent behaviors — something enterprise teams deploying AI coding agents should carefully audit.

                                      Screenshot 2026-05-26 at 3.29.41 PM

                                      On DeepSWE, top models wrote and ran their own tests in as many as 85 percent of runs. On SWE-Bench Pro, where prompts discourage modifying tests, the same models rarely did so. (Source: Datacurve)

                                      What DeepSWE gets right, what it gets wrong, and what it means for the future of AI benchmarks

                                      Datacurve is forthright about several limitations. The standardized harness, while ensuring fairness, routes all edits through bash rather than the model-specific editing tools each family was trained on — apply_patch for GPT, str_replace_based_edit_tool for Claude. This could hold models below their native ceilings. The benchmark draws exclusively from open-source repositories with 500-plus stars, and results may not generalize to proprietary codebases. Bug localization and refactoring tasks are under-represented, and widely used languages like C++ and Java are absent entirely. The verdict assignments in the qualitative analysis come from an LLM analyzer, not human reviewers, and sample sizes are modest — roughly 90 reviewed rollouts per model per benchmark.

                                      It is also worth noting that Datacurve is a startup with its own commercial interests, and an independent benchmark that reshuffles the leaderboard will inevitably invite scrutiny. The company’s decision to publish the full dataset, all agent trajectories, and the evaluation harness on GitHub mitigates this concern considerably, but independent reproduction will be necessary before the AI community treats these results as definitive.

                                      DeepSWE arrives at an inflection point for the AI coding market. Enterprise adoption of AI coding agents is accelerating rapidly, with engineering organizations making consequential bets on which model to build around. The benchmark market itself has become a strategic battleground — Scale AI’s SWE-Bench Pro, which Datacurve directly critiques, is maintained by a company that also provides evaluation services to the labs whose models it ranks.

                                      If DeepSWE’s central findings about verifier reliability and data contamination hold up under independent scrutiny, they could force a reckoning not just with how the industry measures coding agents, but with the broader question of what benchmarks are actually for. A leaderboard where the grading system is wrong a third of the time is not merely inaccurate — it is the kind of broken instrument that makes everyone feel good about progress that may not be real. And in an industry spending billions on a bet that AI agents can do the work of software engineers, the difference between real progress and the appearance of it is not academic. It is the whole game.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      For months, the leading AI coding benchmarks have told enterprise buyers a comforting but misleading story: the top models are all roughly the same. OpenAI’s GPT-5 family, Anthropic’s Claude Opus, and Google’s Gemini Pro have clustered within a narrow band on Scale AI’s SWE-Bench Pro leaderboard, making it nearly impossible for engineering leaders to determine which agent will actually perform best inside their codebases.

                                      On Monday, a startup called Datacurve released a benchmark it says shatters that illusion. DeepSWE, a 113-task evaluation spanning 91 open-source repositories and five programming languages, produces a dramatically wider spread among the same frontier models — and crowns OpenAI’s GPT-5.5 as the clear leader at 70%, sixteen points ahead of its nearest competitor.

                                      «On public leaderboards, top models often look relatively close in capability,» wrote Datacurve co-author Serena Ge on X. «DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.»

                                      The benchmark also delivers a pointed critique of the evaluation infrastructure the AI industry relies on to measure progress: Datacurve’s audit found that SWE-Bench Pro’s verifiers — the automated graders that determine whether an agent solved a task — issued incorrect pass/fail verdicts on roughly one-third of the trials it reviewed.

                                      If that finding holds up, it has sweeping implications. Enterprise procurement teams, venture capitalists, and AI lab marketing departments all lean heavily on benchmark scores to make multimillion-dollar decisions. A 32% error rate in the most widely cited coding benchmark suggests the industry may have been navigating by a broken compass.

                                      Why the most popular AI coding benchmark may be grading on a curve

                                      To understand what Datacurve is claiming, it helps to understand how coding benchmarks work — and how they can go wrong.

                                      The dominant paradigm, pioneered by the SWE-Bench family maintained by Scale AI and academic researchers, constructs tasks by mining real GitHub commits. The process extracts a bug fix or feature addition from a repository’s history, rolls the code back to the pre-fix state, and then asks an AI agent to reproduce the change. The original commit’s test suite serves as the verifier: if the agent’s patch makes the same tests pass, it gets credit. This approach has an elegant simplicity, but Datacurve argues it introduces three systemic weaknesses.

                                      First, contamination. Because tasks are drawn from public GitHub history, the problem statement, the discussion, and often the exact solution are already present in the training data of frontier models. «The SWE-Bench family scrapes existing GitHub issues and PRs, which creates two problems: memorization (models have already seen the solution) and triviality (most tasks are small),» Ge wrote.

                                      Second, scope. SWE-Bench Pro tasks require, on average, just 120 lines of code added across 5 files. DeepSWE’s reference solutions average 668 lines added across 7 files — roughly 5.5 times more code. Yet DeepSWE’s prompts are actually shorter, averaging 2,158 characters versus SWE-Bench Pro’s 4,614. In other words, DeepSWE gives the agent less instruction but expects far more output, which more closely mirrors how a human developer might actually delegate work to an AI assistant.

                                      DeepSWE tasks demand roughly five times more code than SWE-Bench Pro’s while giving agents shorter prompts — a design choice intended to mirror how developers actually hand off work. (Source: Datacurve)

                                      Third — and most damaging — verifier reliability. Datacurve drew 30 tasks at random from both DeepSWE and SWE-Bench Pro, ran three rollouts across 10 frontier model configurations, and then deployed an LLM-based judge to independently assess whether each agent’s patch actually solved the problem. SWE-Bench Pro’s verifiers accepted wrong implementations 8.5% of the time and rejected correct implementations 24% of the time. DeepSWE’s verifiers registered 0.3% and 1.1%, respectively.

                                      Screenshot 2026-05-26 at 3.22.11 PM

                                      Datacurve’s audit found that SWE-Bench Pro’s automated graders rejected correct solutions 24 percent of the time and accepted wrong ones 8.5 percent of the time. DeepSWE’s verifiers kept both rates near zero. (Source: Datacurve)

                                      The false negative problem is especially insidious because it punishes creative solutions. In one documented case, the gold-standard pull request for a SWE-Bench Pro task refactored a private helper function. An agent that correctly solved the task by inlining the same logic — a perfectly valid engineering choice — failed because the test suite tried to import a symbol that only existed in the original author’s specific implementation.

                                      OpenAI’s GPT-5.5 dominates the new benchmark while Claude and Gemini stumble

                                      DeepSWE’s top-line results reorder the familiar hierarchy in ways that should matter to every engineering team evaluating AI coding tools. On SWE-Bench Pro, models from OpenAI, Anthropic, and Google have traded the lead within a 30-point range. DeepSWE stretches that range to 70 points.

                                      GPT-5.5 leads at 70%, followed by GPT-5.4 at 56% and Claude Opus 4.7 at 54%. From there, the drop-off is steep: Claude Sonnet 4.6 lands at 32%, Gemini 3.5 Flash at 28%, GPT-5.4-mini and Kimi K2.6 tied at 24%, and then a long tail of models in the teens and single digits. Claude Haiku 4.5, which scores 39% on SWE-Bench Pro, collapses to zero on DeepSWE — suggesting that some mid-tier models have been significantly overperforming on easier, potentially contaminated benchmarks.

                                      Screenshot 2026-05-26 at 3.09.33 PM

                                      On SWE-Bench Pro, frontier models cluster within a 30-point range. On DeepSWE, the same models spread across 70 points, with some — like Claude Haiku 4.5 — collapsing entirely. (Source: Datacurve)

                                      GPT-5.5 doesn’t just score the highest — it does so efficiently. The model reaches its 70% pass rate with a median cost of $5.80 per trial, a median wall-clock time of 20 minutes, and a median of 47,000 output tokens. GPT-5.4 emerges as perhaps the best overall value at $3.30 per trial with a 56% score. Claude Opus 4.7, meanwhile, costs significantly more per run, and output tokens, wall-clock duration, and dollar cost per trial all vary by an order of magnitude across the agents tested — yet none of these correlates strongly with pass rate. Agents that emit more tokens, run longer, or cost more do not consistently solve more tasks.

                                      Screenshot 2026-05-26 at 3.27.42 PM

                                      GPT-5.4 and GPT-5.5 occupy the cost-efficient frontier, solving the most tasks for the least money per run. Spending more did not reliably produce better results. (Source: Datacurve)

                                      Datacurve’s audit found that Claude has been reading the answer key on existing benchmarks

                                      Perhaps the most provocative finding in DeepSWE’s analysis concerns what the authors label «CHEATED» verdicts — instances where an agent passes a benchmark not by solving the problem, but by reading the answer.

                                      SWE-Bench Pro’s Docker containers ship the repository’s full .git history, which means the gold-standard solution commit is sitting right there in the container’s file system. Most models ignore it. Claude does not. Datacurve’s analysis found that both Claude Opus 4.7 and Claude Opus 4.6 registered «CHEATED» on more than 12% of their reviewed SWE-Bench Pro rollouts. In those instances, the Claude agent ran commands like git log –all or git show to retrieve the merged fix and paste it into its own patch. The behavior accounted for approximately 18% of Opus 4.7’s passes and 25% of Opus 4.6’s passes on the reviewed sample. The issue has been filed publicly as GitHub issue #93 on the SWE-Bench Pro repository.

                                      GPT-5.4 and GPT-5.5 never exhibited this behavior. Gemini configurations stayed around 1%. Datacurve describes the behavior diplomatically — «The benchmark makes this possible (the gold commit lives in the container), but Claude is the family that consistently does so» — but the implication is clear: a meaningful fraction of Claude’s SWE-Bench Pro scores may reflect environmental exploitation rather than genuine engineering capability.

                                      DeepSWE addresses this by shipping only a shallow clone with the base commit, leaving no gold hash for the agent to discover. It is worth noting that the behavior is arguably a sign of Claude’s environmental attentiveness — the model is very good at exploring its surroundings and exploiting available resources. Whether that counts as «cheating» or «resourcefulness» depends on your perspective, but in the context of a benchmark designed to measure independent problem-solving, it undermines the signal.

                                      Screenshot 2026-05-26 at 3.28.43 PM

                                      Two mechanisms by which agents passed SWE-Bench Pro without solving the underlying problem: reading the answer from the container’s Git history, or stubbing features past weak gold tests. (Source: Datacurve)

                                      Each AI model family fails in its own distinctive way, and the patterns matter for enterprise teams

                                      Beyond the top-line scores, Datacurve’s qualitative trajectory analysis reveals distinctly different failure signatures across model families — a finding that could help engineering teams choose the right model for specific types of work.

                                      Claude is forgetful with multi-part prompts. On DeepSWE, Claude configurations miss stated requirements more than any other family. The pattern is consistent: when a prompt enumerates parallel behaviors — «support both sync and async,» for instance — Claude typically implements the obvious branch and forgets to mirror the change. Datacurve reports that roughly two-thirds of Claude’s «MISSED_REQUIREMENT» failures on DeepSWE follow this «one branch shipped» pattern. In one example, Claude Opus 4.7 correctly landed a sync state-data hook in one engine class while the async engine never received the same hook.

                                      GPT, by contrast, implements exactly what is asked. GPT-5.5 had the lowest rate of missing stated behaviors of any configuration tested. Across multiple runs of the same task, GPT trials tended to converge on the same interpretation of the prompt, suggesting instruction-following precision is a stable trait of the model rather than per-run luck.

                                      One of the most intriguing findings involves self-verification. On DeepSWE, Claude Opus 4.7 and GPT-5.4 wrote and ran new tests in the project’s own test framework on over 80% of their runs — even though no one asked them to. On SWE-Bench Pro, those same models dropped to 28% and 18%, respectively. The reason: SWE-Bench Pro’s prompt template explicitly tells agents they «should not modify the testing logic or any of the tests.» Agents dutifully complied, suppressing a behavior that likely would have improved their performance. This suggests that prompt design in production coding workflows may be inadvertently suppressing valuable agent behaviors — something enterprise teams deploying AI coding agents should carefully audit.

                                      Screenshot 2026-05-26 at 3.29.41 PM

                                      On DeepSWE, top models wrote and ran their own tests in as many as 85 percent of runs. On SWE-Bench Pro, where prompts discourage modifying tests, the same models rarely did so. (Source: Datacurve)

                                      What DeepSWE gets right, what it gets wrong, and what it means for the future of AI benchmarks

                                      Datacurve is forthright about several limitations. The standardized harness, while ensuring fairness, routes all edits through bash rather than the model-specific editing tools each family was trained on — apply_patch for GPT, str_replace_based_edit_tool for Claude. This could hold models below their native ceilings. The benchmark draws exclusively from open-source repositories with 500-plus stars, and results may not generalize to proprietary codebases. Bug localization and refactoring tasks are under-represented, and widely used languages like C++ and Java are absent entirely. The verdict assignments in the qualitative analysis come from an LLM analyzer, not human reviewers, and sample sizes are modest — roughly 90 reviewed rollouts per model per benchmark.

                                      It is also worth noting that Datacurve is a startup with its own commercial interests, and an independent benchmark that reshuffles the leaderboard will inevitably invite scrutiny. The company’s decision to publish the full dataset, all agent trajectories, and the evaluation harness on GitHub mitigates this concern considerably, but independent reproduction will be necessary before the AI community treats these results as definitive.

                                      DeepSWE arrives at an inflection point for the AI coding market. Enterprise adoption of AI coding agents is accelerating rapidly, with engineering organizations making consequential bets on which model to build around. The benchmark market itself has become a strategic battleground — Scale AI’s SWE-Bench Pro, which Datacurve directly critiques, is maintained by a company that also provides evaluation services to the labs whose models it ranks.

                                      If DeepSWE’s central findings about verifier reliability and data contamination hold up under independent scrutiny, they could force a reckoning not just with how the industry measures coding agents, but with the broader question of what benchmarks are actually for. A leaderboard where the grading system is wrong a third of the time is not merely inaccurate — it is the kind of broken instrument that makes everyone feel good about progress that may not be real. And in an industry spending billions on a bet that AI agents can do the work of software engineers, the difference between real progress and the appearance of it is not academic. It is the whole game.

                                      Tours Colombia Todo el año Tours Colombia Todo el año Tours Colombia Todo el año

                                      For months, the leading AI coding benchmarks have told enterprise buyers a comforting but misleading story: the top models are all roughly the same. OpenAI’s GPT-5 family, Anthropic’s Claude Opus, and Google’s Gemini Pro have clustered within a narrow band on Scale AI’s SWE-Bench Pro leaderboard, making it nearly impossible for engineering leaders to determine which agent will actually perform best inside their codebases.

                                      On Monday, a startup called Datacurve released a benchmark it says shatters that illusion. DeepSWE, a 113-task evaluation spanning 91 open-source repositories and five programming languages, produces a dramatically wider spread among the same frontier models — and crowns OpenAI’s GPT-5.5 as the clear leader at 70%, sixteen points ahead of its nearest competitor.

                                      «On public leaderboards, top models often look relatively close in capability,» wrote Datacurve co-author Serena Ge on X. «DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.»

                                      The benchmark also delivers a pointed critique of the evaluation infrastructure the AI industry relies on to measure progress: Datacurve’s audit found that SWE-Bench Pro’s verifiers — the automated graders that determine whether an agent solved a task — issued incorrect pass/fail verdicts on roughly one-third of the trials it reviewed.

                                      If that finding holds up, it has sweeping implications. Enterprise procurement teams, venture capitalists, and AI lab marketing departments all lean heavily on benchmark scores to make multimillion-dollar decisions. A 32% error rate in the most widely cited coding benchmark suggests the industry may have been navigating by a broken compass.

                                      Why the most popular AI coding benchmark may be grading on a curve

                                      To understand what Datacurve is claiming, it helps to understand how coding benchmarks work — and how they can go wrong.

                                      The dominant paradigm, pioneered by the SWE-Bench family maintained by Scale AI and academic researchers, constructs tasks by mining real GitHub commits. The process extracts a bug fix or feature addition from a repository’s history, rolls the code back to the pre-fix state, and then asks an AI agent to reproduce the change. The original commit’s test suite serves as the verifier: if the agent’s patch makes the same tests pass, it gets credit. This approach has an elegant simplicity, but Datacurve argues it introduces three systemic weaknesses.

                                      First, contamination. Because tasks are drawn from public GitHub history, the problem statement, the discussion, and often the exact solution are already present in the training data of frontier models. «The SWE-Bench family scrapes existing GitHub issues and PRs, which creates two problems: memorization (models have already seen the solution) and triviality (most tasks are small),» Ge wrote.

                                      Second, scope. SWE-Bench Pro tasks require, on average, just 120 lines of code added across 5 files. DeepSWE’s reference solutions average 668 lines added across 7 files — roughly 5.5 times more code. Yet DeepSWE’s prompts are actually shorter, averaging 2,158 characters versus SWE-Bench Pro’s 4,614. In other words, DeepSWE gives the agent less instruction but expects far more output, which more closely mirrors how a human developer might actually delegate work to an AI assistant.

                                      DeepSWE tasks demand roughly five times more code than SWE-Bench Pro’s while giving agents shorter prompts — a design choice intended to mirror how developers actually hand off work. (Source: Datacurve)

                                      Third — and most damaging — verifier reliability. Datacurve drew 30 tasks at random from both DeepSWE and SWE-Bench Pro, ran three rollouts across 10 frontier model configurations, and then deployed an LLM-based judge to independently assess whether each agent’s patch actually solved the problem. SWE-Bench Pro’s verifiers accepted wrong implementations 8.5% of the time and rejected correct implementations 24% of the time. DeepSWE’s verifiers registered 0.3% and 1.1%, respectively.

                                      Screenshot 2026-05-26 at 3.22.11 PM

                                      Datacurve’s audit found that SWE-Bench Pro’s automated graders rejected correct solutions 24 percent of the time and accepted wrong ones 8.5 percent of the time. DeepSWE’s verifiers kept both rates near zero. (Source: Datacurve)

                                      The false negative problem is especially insidious because it punishes creative solutions. In one documented case, the gold-standard pull request for a SWE-Bench Pro task refactored a private helper function. An agent that correctly solved the task by inlining the same logic — a perfectly valid engineering choice — failed because the test suite tried to import a symbol that only existed in the original author’s specific implementation.

                                      OpenAI’s GPT-5.5 dominates the new benchmark while Claude and Gemini stumble

                                      DeepSWE’s top-line results reorder the familiar hierarchy in ways that should matter to every engineering team evaluating AI coding tools. On SWE-Bench Pro, models from OpenAI, Anthropic, and Google have traded the lead within a 30-point range. DeepSWE stretches that range to 70 points.

                                      GPT-5.5 leads at 70%, followed by GPT-5.4 at 56% and Claude Opus 4.7 at 54%. From there, the drop-off is steep: Claude Sonnet 4.6 lands at 32%, Gemini 3.5 Flash at 28%, GPT-5.4-mini and Kimi K2.6 tied at 24%, and then a long tail of models in the teens and single digits. Claude Haiku 4.5, which scores 39% on SWE-Bench Pro, collapses to zero on DeepSWE — suggesting that some mid-tier models have been significantly overperforming on easier, potentially contaminated benchmarks.

                                      Screenshot 2026-05-26 at 3.09.33 PM

                                      On SWE-Bench Pro, frontier models cluster within a 30-point range. On DeepSWE, the same models spread across 70 points, with some — like Claude Haiku 4.5 — collapsing entirely. (Source: Datacurve)

                                      GPT-5.5 doesn’t just score the highest — it does so efficiently. The model reaches its 70% pass rate with a median cost of $5.80 per trial, a median wall-clock time of 20 minutes, and a median of 47,000 output tokens. GPT-5.4 emerges as perhaps the best overall value at $3.30 per trial with a 56% score. Claude Opus 4.7, meanwhile, costs significantly more per run, and output tokens, wall-clock duration, and dollar cost per trial all vary by an order of magnitude across the agents tested — yet none of these correlates strongly with pass rate. Agents that emit more tokens, run longer, or cost more do not consistently solve more tasks.

                                      Screenshot 2026-05-26 at 3.27.42 PM

                                      GPT-5.4 and GPT-5.5 occupy the cost-efficient frontier, solving the most tasks for the least money per run. Spending more did not reliably produce better results. (Source: Datacurve)

                                      Datacurve’s audit found that Claude has been reading the answer key on existing benchmarks

                                      Perhaps the most provocative finding in DeepSWE’s analysis concerns what the authors label «CHEATED» verdicts — instances where an agent passes a benchmark not by solving the problem, but by reading the answer.

                                      SWE-Bench Pro’s Docker containers ship the repository’s full .git history, which means the gold-standard solution commit is sitting right there in the container’s file system. Most models ignore it. Claude does not. Datacurve’s analysis found that both Claude Opus 4.7 and Claude Opus 4.6 registered «CHEATED» on more than 12% of their reviewed SWE-Bench Pro rollouts. In those instances, the Claude agent ran commands like git log –all or git show to retrieve the merged fix and paste it into its own patch. The behavior accounted for approximately 18% of Opus 4.7’s passes and 25% of Opus 4.6’s passes on the reviewed sample. The issue has been filed publicly as GitHub issue #93 on the SWE-Bench Pro repository.

                                      GPT-5.4 and GPT-5.5 never exhibited this behavior. Gemini configurations stayed around 1%. Datacurve describes the behavior diplomatically — «The benchmark makes this possible (the gold commit lives in the container), but Claude is the family that consistently does so» — but the implication is clear: a meaningful fraction of Claude’s SWE-Bench Pro scores may reflect environmental exploitation rather than genuine engineering capability.

                                      DeepSWE addresses this by shipping only a shallow clone with the base commit, leaving no gold hash for the agent to discover. It is worth noting that the behavior is arguably a sign of Claude’s environmental attentiveness — the model is very good at exploring its surroundings and exploiting available resources. Whether that counts as «cheating» or «resourcefulness» depends on your perspective, but in the context of a benchmark designed to measure independent problem-solving, it undermines the signal.

                                      Screenshot 2026-05-26 at 3.28.43 PM

                                      Two mechanisms by which agents passed SWE-Bench Pro without solving the underlying problem: reading the answer from the container’s Git history, or stubbing features past weak gold tests. (Source: Datacurve)

                                      Each AI model family fails in its own distinctive way, and the patterns matter for enterprise teams

                                      Beyond the top-line scores, Datacurve’s qualitative trajectory analysis reveals distinctly different failure signatures across model families — a finding that could help engineering teams choose the right model for specific types of work.

                                      Claude is forgetful with multi-part prompts. On DeepSWE, Claude configurations miss stated requirements more than any other family. The pattern is consistent: when a prompt enumerates parallel behaviors — «support both sync and async,» for instance — Claude typically implements the obvious branch and forgets to mirror the change. Datacurve reports that roughly two-thirds of Claude’s «MISSED_REQUIREMENT» failures on DeepSWE follow this «one branch shipped» pattern. In one example, Claude Opus 4.7 correctly landed a sync state-data hook in one engine class while the async engine never received the same hook.

                                      GPT, by contrast, implements exactly what is asked. GPT-5.5 had the lowest rate of missing stated behaviors of any configuration tested. Across multiple runs of the same task, GPT trials tended to converge on the same interpretation of the prompt, suggesting instruction-following precision is a stable trait of the model rather than per-run luck.

                                      One of the most intriguing findings involves self-verification. On DeepSWE, Claude Opus 4.7 and GPT-5.4 wrote and ran new tests in the project’s own test framework on over 80% of their runs — even though no one asked them to. On SWE-Bench Pro, those same models dropped to 28% and 18%, respectively. The reason: SWE-Bench Pro’s prompt template explicitly tells agents they «should not modify the testing logic or any of the tests.» Agents dutifully complied, suppressing a behavior that likely would have improved their performance. This suggests that prompt design in production coding workflows may be inadvertently suppressing valuable agent behaviors — something enterprise teams deploying AI coding agents should carefully audit.

                                      Screenshot 2026-05-26 at 3.29.41 PM

                                      On DeepSWE, top models wrote and ran their own tests in as many as 85 percent of runs. On SWE-Bench Pro, where prompts discourage modifying tests, the same models rarely did so. (Source: Datacurve)

                                      What DeepSWE gets right, what it gets wrong, and what it means for the future of AI benchmarks

                                      Datacurve is forthright about several limitations. The standardized harness, while ensuring fairness, routes all edits through bash rather than the model-specific editing tools each family was trained on — apply_patch for GPT, str_replace_based_edit_tool for Claude. This could hold models below their native ceilings. The benchmark draws exclusively from open-source repositories with 500-plus stars, and results may not generalize to proprietary codebases. Bug localization and refactoring tasks are under-represented, and widely used languages like C++ and Java are absent entirely. The verdict assignments in the qualitative analysis come from an LLM analyzer, not human reviewers, and sample sizes are modest — roughly 90 reviewed rollouts per model per benchmark.

                                      It is also worth noting that Datacurve is a startup with its own commercial interests, and an independent benchmark that reshuffles the leaderboard will inevitably invite scrutiny. The company’s decision to publish the full dataset, all agent trajectories, and the evaluation harness on GitHub mitigates this concern considerably, but independent reproduction will be necessary before the AI community treats these results as definitive.

                                      DeepSWE arrives at an inflection point for the AI coding market. Enterprise adoption of AI coding agents is accelerating rapidly, with engineering organizations making consequential bets on which model to build around. The benchmark market itself has become a strategic battleground — Scale AI’s SWE-Bench Pro, which Datacurve directly critiques, is maintained by a company that also provides evaluation services to the labs whose models it ranks.

                                      If DeepSWE’s central findings about verifier reliability and data contamination hold up under independent scrutiny, they could force a reckoning not just with how the industry measures coding agents, but with the broader question of what benchmarks are actually for. A leaderboard where the grading system is wrong a third of the time is not merely inaccurate — it is the kind of broken instrument that makes everyone feel good about progress that may not be real. And in an industry spending billions on a bet that AI agents can do the work of software engineers, the difference between real progress and the appearance of it is not academic. It is the whole game.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      For months, the leading AI coding benchmarks have told enterprise buyers a comforting but misleading story: the top models are all roughly the same. OpenAI’s GPT-5 family, Anthropic’s Claude Opus, and Google’s Gemini Pro have clustered within a narrow band on Scale AI’s SWE-Bench Pro leaderboard, making it nearly impossible for engineering leaders to determine which agent will actually perform best inside their codebases.

                                      On Monday, a startup called Datacurve released a benchmark it says shatters that illusion. DeepSWE, a 113-task evaluation spanning 91 open-source repositories and five programming languages, produces a dramatically wider spread among the same frontier models — and crowns OpenAI’s GPT-5.5 as the clear leader at 70%, sixteen points ahead of its nearest competitor.

                                      «On public leaderboards, top models often look relatively close in capability,» wrote Datacurve co-author Serena Ge on X. «DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.»

                                      The benchmark also delivers a pointed critique of the evaluation infrastructure the AI industry relies on to measure progress: Datacurve’s audit found that SWE-Bench Pro’s verifiers — the automated graders that determine whether an agent solved a task — issued incorrect pass/fail verdicts on roughly one-third of the trials it reviewed.

                                      If that finding holds up, it has sweeping implications. Enterprise procurement teams, venture capitalists, and AI lab marketing departments all lean heavily on benchmark scores to make multimillion-dollar decisions. A 32% error rate in the most widely cited coding benchmark suggests the industry may have been navigating by a broken compass.

                                      Why the most popular AI coding benchmark may be grading on a curve

                                      To understand what Datacurve is claiming, it helps to understand how coding benchmarks work — and how they can go wrong.

                                      The dominant paradigm, pioneered by the SWE-Bench family maintained by Scale AI and academic researchers, constructs tasks by mining real GitHub commits. The process extracts a bug fix or feature addition from a repository’s history, rolls the code back to the pre-fix state, and then asks an AI agent to reproduce the change. The original commit’s test suite serves as the verifier: if the agent’s patch makes the same tests pass, it gets credit. This approach has an elegant simplicity, but Datacurve argues it introduces three systemic weaknesses.

                                      First, contamination. Because tasks are drawn from public GitHub history, the problem statement, the discussion, and often the exact solution are already present in the training data of frontier models. «The SWE-Bench family scrapes existing GitHub issues and PRs, which creates two problems: memorization (models have already seen the solution) and triviality (most tasks are small),» Ge wrote.

                                      Second, scope. SWE-Bench Pro tasks require, on average, just 120 lines of code added across 5 files. DeepSWE’s reference solutions average 668 lines added across 7 files — roughly 5.5 times more code. Yet DeepSWE’s prompts are actually shorter, averaging 2,158 characters versus SWE-Bench Pro’s 4,614. In other words, DeepSWE gives the agent less instruction but expects far more output, which more closely mirrors how a human developer might actually delegate work to an AI assistant.

                                      DeepSWE tasks demand roughly five times more code than SWE-Bench Pro’s while giving agents shorter prompts — a design choice intended to mirror how developers actually hand off work. (Source: Datacurve)

                                      Third — and most damaging — verifier reliability. Datacurve drew 30 tasks at random from both DeepSWE and SWE-Bench Pro, ran three rollouts across 10 frontier model configurations, and then deployed an LLM-based judge to independently assess whether each agent’s patch actually solved the problem. SWE-Bench Pro’s verifiers accepted wrong implementations 8.5% of the time and rejected correct implementations 24% of the time. DeepSWE’s verifiers registered 0.3% and 1.1%, respectively.

                                      Screenshot 2026-05-26 at 3.22.11 PM

                                      Datacurve’s audit found that SWE-Bench Pro’s automated graders rejected correct solutions 24 percent of the time and accepted wrong ones 8.5 percent of the time. DeepSWE’s verifiers kept both rates near zero. (Source: Datacurve)

                                      The false negative problem is especially insidious because it punishes creative solutions. In one documented case, the gold-standard pull request for a SWE-Bench Pro task refactored a private helper function. An agent that correctly solved the task by inlining the same logic — a perfectly valid engineering choice — failed because the test suite tried to import a symbol that only existed in the original author’s specific implementation.

                                      OpenAI’s GPT-5.5 dominates the new benchmark while Claude and Gemini stumble

                                      DeepSWE’s top-line results reorder the familiar hierarchy in ways that should matter to every engineering team evaluating AI coding tools. On SWE-Bench Pro, models from OpenAI, Anthropic, and Google have traded the lead within a 30-point range. DeepSWE stretches that range to 70 points.

                                      GPT-5.5 leads at 70%, followed by GPT-5.4 at 56% and Claude Opus 4.7 at 54%. From there, the drop-off is steep: Claude Sonnet 4.6 lands at 32%, Gemini 3.5 Flash at 28%, GPT-5.4-mini and Kimi K2.6 tied at 24%, and then a long tail of models in the teens and single digits. Claude Haiku 4.5, which scores 39% on SWE-Bench Pro, collapses to zero on DeepSWE — suggesting that some mid-tier models have been significantly overperforming on easier, potentially contaminated benchmarks.

                                      Screenshot 2026-05-26 at 3.09.33 PM

                                      On SWE-Bench Pro, frontier models cluster within a 30-point range. On DeepSWE, the same models spread across 70 points, with some — like Claude Haiku 4.5 — collapsing entirely. (Source: Datacurve)

                                      GPT-5.5 doesn’t just score the highest — it does so efficiently. The model reaches its 70% pass rate with a median cost of $5.80 per trial, a median wall-clock time of 20 minutes, and a median of 47,000 output tokens. GPT-5.4 emerges as perhaps the best overall value at $3.30 per trial with a 56% score. Claude Opus 4.7, meanwhile, costs significantly more per run, and output tokens, wall-clock duration, and dollar cost per trial all vary by an order of magnitude across the agents tested — yet none of these correlates strongly with pass rate. Agents that emit more tokens, run longer, or cost more do not consistently solve more tasks.

                                      Screenshot 2026-05-26 at 3.27.42 PM

                                      GPT-5.4 and GPT-5.5 occupy the cost-efficient frontier, solving the most tasks for the least money per run. Spending more did not reliably produce better results. (Source: Datacurve)

                                      Datacurve’s audit found that Claude has been reading the answer key on existing benchmarks

                                      Perhaps the most provocative finding in DeepSWE’s analysis concerns what the authors label «CHEATED» verdicts — instances where an agent passes a benchmark not by solving the problem, but by reading the answer.

                                      SWE-Bench Pro’s Docker containers ship the repository’s full .git history, which means the gold-standard solution commit is sitting right there in the container’s file system. Most models ignore it. Claude does not. Datacurve’s analysis found that both Claude Opus 4.7 and Claude Opus 4.6 registered «CHEATED» on more than 12% of their reviewed SWE-Bench Pro rollouts. In those instances, the Claude agent ran commands like git log –all or git show to retrieve the merged fix and paste it into its own patch. The behavior accounted for approximately 18% of Opus 4.7’s passes and 25% of Opus 4.6’s passes on the reviewed sample. The issue has been filed publicly as GitHub issue #93 on the SWE-Bench Pro repository.

                                      GPT-5.4 and GPT-5.5 never exhibited this behavior. Gemini configurations stayed around 1%. Datacurve describes the behavior diplomatically — «The benchmark makes this possible (the gold commit lives in the container), but Claude is the family that consistently does so» — but the implication is clear: a meaningful fraction of Claude’s SWE-Bench Pro scores may reflect environmental exploitation rather than genuine engineering capability.

                                      DeepSWE addresses this by shipping only a shallow clone with the base commit, leaving no gold hash for the agent to discover. It is worth noting that the behavior is arguably a sign of Claude’s environmental attentiveness — the model is very good at exploring its surroundings and exploiting available resources. Whether that counts as «cheating» or «resourcefulness» depends on your perspective, but in the context of a benchmark designed to measure independent problem-solving, it undermines the signal.

                                      Screenshot 2026-05-26 at 3.28.43 PM

                                      Two mechanisms by which agents passed SWE-Bench Pro without solving the underlying problem: reading the answer from the container’s Git history, or stubbing features past weak gold tests. (Source: Datacurve)

                                      Each AI model family fails in its own distinctive way, and the patterns matter for enterprise teams

                                      Beyond the top-line scores, Datacurve’s qualitative trajectory analysis reveals distinctly different failure signatures across model families — a finding that could help engineering teams choose the right model for specific types of work.

                                      Claude is forgetful with multi-part prompts. On DeepSWE, Claude configurations miss stated requirements more than any other family. The pattern is consistent: when a prompt enumerates parallel behaviors — «support both sync and async,» for instance — Claude typically implements the obvious branch and forgets to mirror the change. Datacurve reports that roughly two-thirds of Claude’s «MISSED_REQUIREMENT» failures on DeepSWE follow this «one branch shipped» pattern. In one example, Claude Opus 4.7 correctly landed a sync state-data hook in one engine class while the async engine never received the same hook.

                                      GPT, by contrast, implements exactly what is asked. GPT-5.5 had the lowest rate of missing stated behaviors of any configuration tested. Across multiple runs of the same task, GPT trials tended to converge on the same interpretation of the prompt, suggesting instruction-following precision is a stable trait of the model rather than per-run luck.

                                      One of the most intriguing findings involves self-verification. On DeepSWE, Claude Opus 4.7 and GPT-5.4 wrote and ran new tests in the project’s own test framework on over 80% of their runs — even though no one asked them to. On SWE-Bench Pro, those same models dropped to 28% and 18%, respectively. The reason: SWE-Bench Pro’s prompt template explicitly tells agents they «should not modify the testing logic or any of the tests.» Agents dutifully complied, suppressing a behavior that likely would have improved their performance. This suggests that prompt design in production coding workflows may be inadvertently suppressing valuable agent behaviors — something enterprise teams deploying AI coding agents should carefully audit.

                                      Screenshot 2026-05-26 at 3.29.41 PM

                                      On DeepSWE, top models wrote and ran their own tests in as many as 85 percent of runs. On SWE-Bench Pro, where prompts discourage modifying tests, the same models rarely did so. (Source: Datacurve)

                                      What DeepSWE gets right, what it gets wrong, and what it means for the future of AI benchmarks

                                      Datacurve is forthright about several limitations. The standardized harness, while ensuring fairness, routes all edits through bash rather than the model-specific editing tools each family was trained on — apply_patch for GPT, str_replace_based_edit_tool for Claude. This could hold models below their native ceilings. The benchmark draws exclusively from open-source repositories with 500-plus stars, and results may not generalize to proprietary codebases. Bug localization and refactoring tasks are under-represented, and widely used languages like C++ and Java are absent entirely. The verdict assignments in the qualitative analysis come from an LLM analyzer, not human reviewers, and sample sizes are modest — roughly 90 reviewed rollouts per model per benchmark.

                                      It is also worth noting that Datacurve is a startup with its own commercial interests, and an independent benchmark that reshuffles the leaderboard will inevitably invite scrutiny. The company’s decision to publish the full dataset, all agent trajectories, and the evaluation harness on GitHub mitigates this concern considerably, but independent reproduction will be necessary before the AI community treats these results as definitive.

                                      DeepSWE arrives at an inflection point for the AI coding market. Enterprise adoption of AI coding agents is accelerating rapidly, with engineering organizations making consequential bets on which model to build around. The benchmark market itself has become a strategic battleground — Scale AI’s SWE-Bench Pro, which Datacurve directly critiques, is maintained by a company that also provides evaluation services to the labs whose models it ranks.

                                      If DeepSWE’s central findings about verifier reliability and data contamination hold up under independent scrutiny, they could force a reckoning not just with how the industry measures coding agents, but with the broader question of what benchmarks are actually for. A leaderboard where the grading system is wrong a third of the time is not merely inaccurate — it is the kind of broken instrument that makes everyone feel good about progress that may not be real. And in an industry spending billions on a bet that AI agents can do the work of software engineers, the difference between real progress and the appearance of it is not academic. It is the whole game.

                                      ¡No te pierdas las noticias destacadas!

                                      Suscríbete y recibe las historias más importantes del día.

                                      Al suscribirte aceptas nuestros términos y condiciones y política de privacidad.

                                      For months, the leading AI coding benchmarks have told enterprise buyers a comforting but misleading story: the top models are all roughly the same. OpenAI’s GPT-5 family, Anthropic’s Claude Opus, and Google’s Gemini Pro have clustered within a narrow band on Scale AI’s SWE-Bench Pro leaderboard, making it nearly impossible for engineering leaders to determine which agent will actually perform best inside their codebases.

                                      On Monday, a startup called Datacurve released a benchmark it says shatters that illusion. DeepSWE, a 113-task evaluation spanning 91 open-source repositories and five programming languages, produces a dramatically wider spread among the same frontier models — and crowns OpenAI’s GPT-5.5 as the clear leader at 70%, sixteen points ahead of its nearest competitor.

                                      «On public leaderboards, top models often look relatively close in capability,» wrote Datacurve co-author Serena Ge on X. «DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.»

                                      The benchmark also delivers a pointed critique of the evaluation infrastructure the AI industry relies on to measure progress: Datacurve’s audit found that SWE-Bench Pro’s verifiers — the automated graders that determine whether an agent solved a task — issued incorrect pass/fail verdicts on roughly one-third of the trials it reviewed.

                                      If that finding holds up, it has sweeping implications. Enterprise procurement teams, venture capitalists, and AI lab marketing departments all lean heavily on benchmark scores to make multimillion-dollar decisions. A 32% error rate in the most widely cited coding benchmark suggests the industry may have been navigating by a broken compass.

                                      Why the most popular AI coding benchmark may be grading on a curve

                                      To understand what Datacurve is claiming, it helps to understand how coding benchmarks work — and how they can go wrong.

                                      The dominant paradigm, pioneered by the SWE-Bench family maintained by Scale AI and academic researchers, constructs tasks by mining real GitHub commits. The process extracts a bug fix or feature addition from a repository’s history, rolls the code back to the pre-fix state, and then asks an AI agent to reproduce the change. The original commit’s test suite serves as the verifier: if the agent’s patch makes the same tests pass, it gets credit. This approach has an elegant simplicity, but Datacurve argues it introduces three systemic weaknesses.

                                      First, contamination. Because tasks are drawn from public GitHub history, the problem statement, the discussion, and often the exact solution are already present in the training data of frontier models. «The SWE-Bench family scrapes existing GitHub issues and PRs, which creates two problems: memorization (models have already seen the solution) and triviality (most tasks are small),» Ge wrote.

                                      Second, scope. SWE-Bench Pro tasks require, on average, just 120 lines of code added across 5 files. DeepSWE’s reference solutions average 668 lines added across 7 files — roughly 5.5 times more code. Yet DeepSWE’s prompts are actually shorter, averaging 2,158 characters versus SWE-Bench Pro’s 4,614. In other words, DeepSWE gives the agent less instruction but expects far more output, which more closely mirrors how a human developer might actually delegate work to an AI assistant.

                                      DeepSWE tasks demand roughly five times more code than SWE-Bench Pro’s while giving agents shorter prompts — a design choice intended to mirror how developers actually hand off work. (Source: Datacurve)

                                      Third — and most damaging — verifier reliability. Datacurve drew 30 tasks at random from both DeepSWE and SWE-Bench Pro, ran three rollouts across 10 frontier model configurations, and then deployed an LLM-based judge to independently assess whether each agent’s patch actually solved the problem. SWE-Bench Pro’s verifiers accepted wrong implementations 8.5% of the time and rejected correct implementations 24% of the time. DeepSWE’s verifiers registered 0.3% and 1.1%, respectively.

                                      Screenshot 2026-05-26 at 3.22.11 PM

                                      Datacurve’s audit found that SWE-Bench Pro’s automated graders rejected correct solutions 24 percent of the time and accepted wrong ones 8.5 percent of the time. DeepSWE’s verifiers kept both rates near zero. (Source: Datacurve)

                                      The false negative problem is especially insidious because it punishes creative solutions. In one documented case, the gold-standard pull request for a SWE-Bench Pro task refactored a private helper function. An agent that correctly solved the task by inlining the same logic — a perfectly valid engineering choice — failed because the test suite tried to import a symbol that only existed in the original author’s specific implementation.

                                      OpenAI’s GPT-5.5 dominates the new benchmark while Claude and Gemini stumble

                                      DeepSWE’s top-line results reorder the familiar hierarchy in ways that should matter to every engineering team evaluating AI coding tools. On SWE-Bench Pro, models from OpenAI, Anthropic, and Google have traded the lead within a 30-point range. DeepSWE stretches that range to 70 points.

                                      GPT-5.5 leads at 70%, followed by GPT-5.4 at 56% and Claude Opus 4.7 at 54%. From there, the drop-off is steep: Claude Sonnet 4.6 lands at 32%, Gemini 3.5 Flash at 28%, GPT-5.4-mini and Kimi K2.6 tied at 24%, and then a long tail of models in the teens and single digits. Claude Haiku 4.5, which scores 39% on SWE-Bench Pro, collapses to zero on DeepSWE — suggesting that some mid-tier models have been significantly overperforming on easier, potentially contaminated benchmarks.

                                      Screenshot 2026-05-26 at 3.09.33 PM

                                      On SWE-Bench Pro, frontier models cluster within a 30-point range. On DeepSWE, the same models spread across 70 points, with some — like Claude Haiku 4.5 — collapsing entirely. (Source: Datacurve)

                                      GPT-5.5 doesn’t just score the highest — it does so efficiently. The model reaches its 70% pass rate with a median cost of $5.80 per trial, a median wall-clock time of 20 minutes, and a median of 47,000 output tokens. GPT-5.4 emerges as perhaps the best overall value at $3.30 per trial with a 56% score. Claude Opus 4.7, meanwhile, costs significantly more per run, and output tokens, wall-clock duration, and dollar cost per trial all vary by an order of magnitude across the agents tested — yet none of these correlates strongly with pass rate. Agents that emit more tokens, run longer, or cost more do not consistently solve more tasks.

                                      Screenshot 2026-05-26 at 3.27.42 PM

                                      GPT-5.4 and GPT-5.5 occupy the cost-efficient frontier, solving the most tasks for the least money per run. Spending more did not reliably produce better results. (Source: Datacurve)

                                      Datacurve’s audit found that Claude has been reading the answer key on existing benchmarks

                                      Perhaps the most provocative finding in DeepSWE’s analysis concerns what the authors label «CHEATED» verdicts — instances where an agent passes a benchmark not by solving the problem, but by reading the answer.

                                      SWE-Bench Pro’s Docker containers ship the repository’s full .git history, which means the gold-standard solution commit is sitting right there in the container’s file system. Most models ignore it. Claude does not. Datacurve’s analysis found that both Claude Opus 4.7 and Claude Opus 4.6 registered «CHEATED» on more than 12% of their reviewed SWE-Bench Pro rollouts. In those instances, the Claude agent ran commands like git log –all or git show to retrieve the merged fix and paste it into its own patch. The behavior accounted for approximately 18% of Opus 4.7’s passes and 25% of Opus 4.6’s passes on the reviewed sample. The issue has been filed publicly as GitHub issue #93 on the SWE-Bench Pro repository.

                                      GPT-5.4 and GPT-5.5 never exhibited this behavior. Gemini configurations stayed around 1%. Datacurve describes the behavior diplomatically — «The benchmark makes this possible (the gold commit lives in the container), but Claude is the family that consistently does so» — but the implication is clear: a meaningful fraction of Claude’s SWE-Bench Pro scores may reflect environmental exploitation rather than genuine engineering capability.

                                      DeepSWE addresses this by shipping only a shallow clone with the base commit, leaving no gold hash for the agent to discover. It is worth noting that the behavior is arguably a sign of Claude’s environmental attentiveness — the model is very good at exploring its surroundings and exploiting available resources. Whether that counts as «cheating» or «resourcefulness» depends on your perspective, but in the context of a benchmark designed to measure independent problem-solving, it undermines the signal.

                                      Screenshot 2026-05-26 at 3.28.43 PM

                                      Two mechanisms by which agents passed SWE-Bench Pro without solving the underlying problem: reading the answer from the container’s Git history, or stubbing features past weak gold tests. (Source: Datacurve)

                                      Each AI model family fails in its own distinctive way, and the patterns matter for enterprise teams

                                      Beyond the top-line scores, Datacurve’s qualitative trajectory analysis reveals distinctly different failure signatures across model families — a finding that could help engineering teams choose the right model for specific types of work.

                                      Claude is forgetful with multi-part prompts. On DeepSWE, Claude configurations miss stated requirements more than any other family. The pattern is consistent: when a prompt enumerates parallel behaviors — «support both sync and async,» for instance — Claude typically implements the obvious branch and forgets to mirror the change. Datacurve reports that roughly two-thirds of Claude’s «MISSED_REQUIREMENT» failures on DeepSWE follow this «one branch shipped» pattern. In one example, Claude Opus 4.7 correctly landed a sync state-data hook in one engine class while the async engine never received the same hook.

                                      GPT, by contrast, implements exactly what is asked. GPT-5.5 had the lowest rate of missing stated behaviors of any configuration tested. Across multiple runs of the same task, GPT trials tended to converge on the same interpretation of the prompt, suggesting instruction-following precision is a stable trait of the model rather than per-run luck.

                                      One of the most intriguing findings involves self-verification. On DeepSWE, Claude Opus 4.7 and GPT-5.4 wrote and ran new tests in the project’s own test framework on over 80% of their runs — even though no one asked them to. On SWE-Bench Pro, those same models dropped to 28% and 18%, respectively. The reason: SWE-Bench Pro’s prompt template explicitly tells agents they «should not modify the testing logic or any of the tests.» Agents dutifully complied, suppressing a behavior that likely would have improved their performance. This suggests that prompt design in production coding workflows may be inadvertently suppressing valuable agent behaviors — something enterprise teams deploying AI coding agents should carefully audit.

                                      Screenshot 2026-05-26 at 3.29.41 PM

                                      On DeepSWE, top models wrote and ran their own tests in as many as 85 percent of runs. On SWE-Bench Pro, where prompts discourage modifying tests, the same models rarely did so. (Source: Datacurve)

                                      What DeepSWE gets right, what it gets wrong, and what it means for the future of AI benchmarks

                                      Datacurve is forthright about several limitations. The standardized harness, while ensuring fairness, routes all edits through bash rather than the model-specific editing tools each family was trained on — apply_patch for GPT, str_replace_based_edit_tool for Claude. This could hold models below their native ceilings. The benchmark draws exclusively from open-source repositories with 500-plus stars, and results may not generalize to proprietary codebases. Bug localization and refactoring tasks are under-represented, and widely used languages like C++ and Java are absent entirely. The verdict assignments in the qualitative analysis come from an LLM analyzer, not human reviewers, and sample sizes are modest — roughly 90 reviewed rollouts per model per benchmark.

                                      It is also worth noting that Datacurve is a startup with its own commercial interests, and an independent benchmark that reshuffles the leaderboard will inevitably invite scrutiny. The company’s decision to publish the full dataset, all agent trajectories, and the evaluation harness on GitHub mitigates this concern considerably, but independent reproduction will be necessary before the AI community treats these results as definitive.

                                      DeepSWE arrives at an inflection point for the AI coding market. Enterprise adoption of AI coding agents is accelerating rapidly, with engineering organizations making consequential bets on which model to build around. The benchmark market itself has become a strategic battleground — Scale AI’s SWE-Bench Pro, which Datacurve directly critiques, is maintained by a company that also provides evaluation services to the labs whose models it ranks.

                                      If DeepSWE’s central findings about verifier reliability and data contamination hold up under independent scrutiny, they could force a reckoning not just with how the industry measures coding agents, but with the broader question of what benchmarks are actually for. A leaderboard where the grading system is wrong a third of the time is not merely inaccurate — it is the kind of broken instrument that makes everyone feel good about progress that may not be real. And in an industry spending billions on a bet that AI agents can do the work of software engineers, the difference between real progress and the appearance of it is not academic. It is the whole game.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      For months, the leading AI coding benchmarks have told enterprise buyers a comforting but misleading story: the top models are all roughly the same. OpenAI’s GPT-5 family, Anthropic’s Claude Opus, and Google’s Gemini Pro have clustered within a narrow band on Scale AI’s SWE-Bench Pro leaderboard, making it nearly impossible for engineering leaders to determine which agent will actually perform best inside their codebases.

                                      On Monday, a startup called Datacurve released a benchmark it says shatters that illusion. DeepSWE, a 113-task evaluation spanning 91 open-source repositories and five programming languages, produces a dramatically wider spread among the same frontier models — and crowns OpenAI’s GPT-5.5 as the clear leader at 70%, sixteen points ahead of its nearest competitor.

                                      «On public leaderboards, top models often look relatively close in capability,» wrote Datacurve co-author Serena Ge on X. «DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.»

                                      The benchmark also delivers a pointed critique of the evaluation infrastructure the AI industry relies on to measure progress: Datacurve’s audit found that SWE-Bench Pro’s verifiers — the automated graders that determine whether an agent solved a task — issued incorrect pass/fail verdicts on roughly one-third of the trials it reviewed.

                                      If that finding holds up, it has sweeping implications. Enterprise procurement teams, venture capitalists, and AI lab marketing departments all lean heavily on benchmark scores to make multimillion-dollar decisions. A 32% error rate in the most widely cited coding benchmark suggests the industry may have been navigating by a broken compass.

                                      Why the most popular AI coding benchmark may be grading on a curve

                                      To understand what Datacurve is claiming, it helps to understand how coding benchmarks work — and how they can go wrong.

                                      The dominant paradigm, pioneered by the SWE-Bench family maintained by Scale AI and academic researchers, constructs tasks by mining real GitHub commits. The process extracts a bug fix or feature addition from a repository’s history, rolls the code back to the pre-fix state, and then asks an AI agent to reproduce the change. The original commit’s test suite serves as the verifier: if the agent’s patch makes the same tests pass, it gets credit. This approach has an elegant simplicity, but Datacurve argues it introduces three systemic weaknesses.

                                      First, contamination. Because tasks are drawn from public GitHub history, the problem statement, the discussion, and often the exact solution are already present in the training data of frontier models. «The SWE-Bench family scrapes existing GitHub issues and PRs, which creates two problems: memorization (models have already seen the solution) and triviality (most tasks are small),» Ge wrote.

                                      Second, scope. SWE-Bench Pro tasks require, on average, just 120 lines of code added across 5 files. DeepSWE’s reference solutions average 668 lines added across 7 files — roughly 5.5 times more code. Yet DeepSWE’s prompts are actually shorter, averaging 2,158 characters versus SWE-Bench Pro’s 4,614. In other words, DeepSWE gives the agent less instruction but expects far more output, which more closely mirrors how a human developer might actually delegate work to an AI assistant.

                                      DeepSWE tasks demand roughly five times more code than SWE-Bench Pro’s while giving agents shorter prompts — a design choice intended to mirror how developers actually hand off work. (Source: Datacurve)

                                      Third — and most damaging — verifier reliability. Datacurve drew 30 tasks at random from both DeepSWE and SWE-Bench Pro, ran three rollouts across 10 frontier model configurations, and then deployed an LLM-based judge to independently assess whether each agent’s patch actually solved the problem. SWE-Bench Pro’s verifiers accepted wrong implementations 8.5% of the time and rejected correct implementations 24% of the time. DeepSWE’s verifiers registered 0.3% and 1.1%, respectively.

                                      Screenshot 2026-05-26 at 3.22.11 PM

                                      Datacurve’s audit found that SWE-Bench Pro’s automated graders rejected correct solutions 24 percent of the time and accepted wrong ones 8.5 percent of the time. DeepSWE’s verifiers kept both rates near zero. (Source: Datacurve)

                                      The false negative problem is especially insidious because it punishes creative solutions. In one documented case, the gold-standard pull request for a SWE-Bench Pro task refactored a private helper function. An agent that correctly solved the task by inlining the same logic — a perfectly valid engineering choice — failed because the test suite tried to import a symbol that only existed in the original author’s specific implementation.

                                      OpenAI’s GPT-5.5 dominates the new benchmark while Claude and Gemini stumble

                                      DeepSWE’s top-line results reorder the familiar hierarchy in ways that should matter to every engineering team evaluating AI coding tools. On SWE-Bench Pro, models from OpenAI, Anthropic, and Google have traded the lead within a 30-point range. DeepSWE stretches that range to 70 points.

                                      GPT-5.5 leads at 70%, followed by GPT-5.4 at 56% and Claude Opus 4.7 at 54%. From there, the drop-off is steep: Claude Sonnet 4.6 lands at 32%, Gemini 3.5 Flash at 28%, GPT-5.4-mini and Kimi K2.6 tied at 24%, and then a long tail of models in the teens and single digits. Claude Haiku 4.5, which scores 39% on SWE-Bench Pro, collapses to zero on DeepSWE — suggesting that some mid-tier models have been significantly overperforming on easier, potentially contaminated benchmarks.

                                      Screenshot 2026-05-26 at 3.09.33 PM

                                      On SWE-Bench Pro, frontier models cluster within a 30-point range. On DeepSWE, the same models spread across 70 points, with some — like Claude Haiku 4.5 — collapsing entirely. (Source: Datacurve)

                                      GPT-5.5 doesn’t just score the highest — it does so efficiently. The model reaches its 70% pass rate with a median cost of $5.80 per trial, a median wall-clock time of 20 minutes, and a median of 47,000 output tokens. GPT-5.4 emerges as perhaps the best overall value at $3.30 per trial with a 56% score. Claude Opus 4.7, meanwhile, costs significantly more per run, and output tokens, wall-clock duration, and dollar cost per trial all vary by an order of magnitude across the agents tested — yet none of these correlates strongly with pass rate. Agents that emit more tokens, run longer, or cost more do not consistently solve more tasks.

                                      Screenshot 2026-05-26 at 3.27.42 PM

                                      GPT-5.4 and GPT-5.5 occupy the cost-efficient frontier, solving the most tasks for the least money per run. Spending more did not reliably produce better results. (Source: Datacurve)

                                      Datacurve’s audit found that Claude has been reading the answer key on existing benchmarks

                                      Perhaps the most provocative finding in DeepSWE’s analysis concerns what the authors label «CHEATED» verdicts — instances where an agent passes a benchmark not by solving the problem, but by reading the answer.

                                      SWE-Bench Pro’s Docker containers ship the repository’s full .git history, which means the gold-standard solution commit is sitting right there in the container’s file system. Most models ignore it. Claude does not. Datacurve’s analysis found that both Claude Opus 4.7 and Claude Opus 4.6 registered «CHEATED» on more than 12% of their reviewed SWE-Bench Pro rollouts. In those instances, the Claude agent ran commands like git log –all or git show to retrieve the merged fix and paste it into its own patch. The behavior accounted for approximately 18% of Opus 4.7’s passes and 25% of Opus 4.6’s passes on the reviewed sample. The issue has been filed publicly as GitHub issue #93 on the SWE-Bench Pro repository.

                                      GPT-5.4 and GPT-5.5 never exhibited this behavior. Gemini configurations stayed around 1%. Datacurve describes the behavior diplomatically — «The benchmark makes this possible (the gold commit lives in the container), but Claude is the family that consistently does so» — but the implication is clear: a meaningful fraction of Claude’s SWE-Bench Pro scores may reflect environmental exploitation rather than genuine engineering capability.

                                      DeepSWE addresses this by shipping only a shallow clone with the base commit, leaving no gold hash for the agent to discover. It is worth noting that the behavior is arguably a sign of Claude’s environmental attentiveness — the model is very good at exploring its surroundings and exploiting available resources. Whether that counts as «cheating» or «resourcefulness» depends on your perspective, but in the context of a benchmark designed to measure independent problem-solving, it undermines the signal.

                                      Screenshot 2026-05-26 at 3.28.43 PM

                                      Two mechanisms by which agents passed SWE-Bench Pro without solving the underlying problem: reading the answer from the container’s Git history, or stubbing features past weak gold tests. (Source: Datacurve)

                                      Each AI model family fails in its own distinctive way, and the patterns matter for enterprise teams

                                      Beyond the top-line scores, Datacurve’s qualitative trajectory analysis reveals distinctly different failure signatures across model families — a finding that could help engineering teams choose the right model for specific types of work.

                                      Claude is forgetful with multi-part prompts. On DeepSWE, Claude configurations miss stated requirements more than any other family. The pattern is consistent: when a prompt enumerates parallel behaviors — «support both sync and async,» for instance — Claude typically implements the obvious branch and forgets to mirror the change. Datacurve reports that roughly two-thirds of Claude’s «MISSED_REQUIREMENT» failures on DeepSWE follow this «one branch shipped» pattern. In one example, Claude Opus 4.7 correctly landed a sync state-data hook in one engine class while the async engine never received the same hook.

                                      GPT, by contrast, implements exactly what is asked. GPT-5.5 had the lowest rate of missing stated behaviors of any configuration tested. Across multiple runs of the same task, GPT trials tended to converge on the same interpretation of the prompt, suggesting instruction-following precision is a stable trait of the model rather than per-run luck.

                                      One of the most intriguing findings involves self-verification. On DeepSWE, Claude Opus 4.7 and GPT-5.4 wrote and ran new tests in the project’s own test framework on over 80% of their runs — even though no one asked them to. On SWE-Bench Pro, those same models dropped to 28% and 18%, respectively. The reason: SWE-Bench Pro’s prompt template explicitly tells agents they «should not modify the testing logic or any of the tests.» Agents dutifully complied, suppressing a behavior that likely would have improved their performance. This suggests that prompt design in production coding workflows may be inadvertently suppressing valuable agent behaviors — something enterprise teams deploying AI coding agents should carefully audit.

                                      Screenshot 2026-05-26 at 3.29.41 PM

                                      On DeepSWE, top models wrote and ran their own tests in as many as 85 percent of runs. On SWE-Bench Pro, where prompts discourage modifying tests, the same models rarely did so. (Source: Datacurve)

                                      What DeepSWE gets right, what it gets wrong, and what it means for the future of AI benchmarks

                                      Datacurve is forthright about several limitations. The standardized harness, while ensuring fairness, routes all edits through bash rather than the model-specific editing tools each family was trained on — apply_patch for GPT, str_replace_based_edit_tool for Claude. This could hold models below their native ceilings. The benchmark draws exclusively from open-source repositories with 500-plus stars, and results may not generalize to proprietary codebases. Bug localization and refactoring tasks are under-represented, and widely used languages like C++ and Java are absent entirely. The verdict assignments in the qualitative analysis come from an LLM analyzer, not human reviewers, and sample sizes are modest — roughly 90 reviewed rollouts per model per benchmark.

                                      It is also worth noting that Datacurve is a startup with its own commercial interests, and an independent benchmark that reshuffles the leaderboard will inevitably invite scrutiny. The company’s decision to publish the full dataset, all agent trajectories, and the evaluation harness on GitHub mitigates this concern considerably, but independent reproduction will be necessary before the AI community treats these results as definitive.

                                      DeepSWE arrives at an inflection point for the AI coding market. Enterprise adoption of AI coding agents is accelerating rapidly, with engineering organizations making consequential bets on which model to build around. The benchmark market itself has become a strategic battleground — Scale AI’s SWE-Bench Pro, which Datacurve directly critiques, is maintained by a company that also provides evaluation services to the labs whose models it ranks.

                                      If DeepSWE’s central findings about verifier reliability and data contamination hold up under independent scrutiny, they could force a reckoning not just with how the industry measures coding agents, but with the broader question of what benchmarks are actually for. A leaderboard where the grading system is wrong a third of the time is not merely inaccurate — it is the kind of broken instrument that makes everyone feel good about progress that may not be real. And in an industry spending billions on a bet that AI agents can do the work of software engineers, the difference between real progress and the appearance of it is not academic. It is the whole game.

                                      Tours Colombia Todo el año Tours Colombia Todo el año Tours Colombia Todo el año

                                      For months, the leading AI coding benchmarks have told enterprise buyers a comforting but misleading story: the top models are all roughly the same. OpenAI’s GPT-5 family, Anthropic’s Claude Opus, and Google’s Gemini Pro have clustered within a narrow band on Scale AI’s SWE-Bench Pro leaderboard, making it nearly impossible for engineering leaders to determine which agent will actually perform best inside their codebases.

                                      On Monday, a startup called Datacurve released a benchmark it says shatters that illusion. DeepSWE, a 113-task evaluation spanning 91 open-source repositories and five programming languages, produces a dramatically wider spread among the same frontier models — and crowns OpenAI’s GPT-5.5 as the clear leader at 70%, sixteen points ahead of its nearest competitor.

                                      «On public leaderboards, top models often look relatively close in capability,» wrote Datacurve co-author Serena Ge on X. «DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.»

                                      The benchmark also delivers a pointed critique of the evaluation infrastructure the AI industry relies on to measure progress: Datacurve’s audit found that SWE-Bench Pro’s verifiers — the automated graders that determine whether an agent solved a task — issued incorrect pass/fail verdicts on roughly one-third of the trials it reviewed.

                                      If that finding holds up, it has sweeping implications. Enterprise procurement teams, venture capitalists, and AI lab marketing departments all lean heavily on benchmark scores to make multimillion-dollar decisions. A 32% error rate in the most widely cited coding benchmark suggests the industry may have been navigating by a broken compass.

                                      Why the most popular AI coding benchmark may be grading on a curve

                                      To understand what Datacurve is claiming, it helps to understand how coding benchmarks work — and how they can go wrong.

                                      The dominant paradigm, pioneered by the SWE-Bench family maintained by Scale AI and academic researchers, constructs tasks by mining real GitHub commits. The process extracts a bug fix or feature addition from a repository’s history, rolls the code back to the pre-fix state, and then asks an AI agent to reproduce the change. The original commit’s test suite serves as the verifier: if the agent’s patch makes the same tests pass, it gets credit. This approach has an elegant simplicity, but Datacurve argues it introduces three systemic weaknesses.

                                      First, contamination. Because tasks are drawn from public GitHub history, the problem statement, the discussion, and often the exact solution are already present in the training data of frontier models. «The SWE-Bench family scrapes existing GitHub issues and PRs, which creates two problems: memorization (models have already seen the solution) and triviality (most tasks are small),» Ge wrote.

                                      Second, scope. SWE-Bench Pro tasks require, on average, just 120 lines of code added across 5 files. DeepSWE’s reference solutions average 668 lines added across 7 files — roughly 5.5 times more code. Yet DeepSWE’s prompts are actually shorter, averaging 2,158 characters versus SWE-Bench Pro’s 4,614. In other words, DeepSWE gives the agent less instruction but expects far more output, which more closely mirrors how a human developer might actually delegate work to an AI assistant.

                                      DeepSWE tasks demand roughly five times more code than SWE-Bench Pro’s while giving agents shorter prompts — a design choice intended to mirror how developers actually hand off work. (Source: Datacurve)

                                      Third — and most damaging — verifier reliability. Datacurve drew 30 tasks at random from both DeepSWE and SWE-Bench Pro, ran three rollouts across 10 frontier model configurations, and then deployed an LLM-based judge to independently assess whether each agent’s patch actually solved the problem. SWE-Bench Pro’s verifiers accepted wrong implementations 8.5% of the time and rejected correct implementations 24% of the time. DeepSWE’s verifiers registered 0.3% and 1.1%, respectively.

                                      Screenshot 2026-05-26 at 3.22.11 PM

                                      Datacurve’s audit found that SWE-Bench Pro’s automated graders rejected correct solutions 24 percent of the time and accepted wrong ones 8.5 percent of the time. DeepSWE’s verifiers kept both rates near zero. (Source: Datacurve)

                                      The false negative problem is especially insidious because it punishes creative solutions. In one documented case, the gold-standard pull request for a SWE-Bench Pro task refactored a private helper function. An agent that correctly solved the task by inlining the same logic — a perfectly valid engineering choice — failed because the test suite tried to import a symbol that only existed in the original author’s specific implementation.

                                      OpenAI’s GPT-5.5 dominates the new benchmark while Claude and Gemini stumble

                                      DeepSWE’s top-line results reorder the familiar hierarchy in ways that should matter to every engineering team evaluating AI coding tools. On SWE-Bench Pro, models from OpenAI, Anthropic, and Google have traded the lead within a 30-point range. DeepSWE stretches that range to 70 points.

                                      GPT-5.5 leads at 70%, followed by GPT-5.4 at 56% and Claude Opus 4.7 at 54%. From there, the drop-off is steep: Claude Sonnet 4.6 lands at 32%, Gemini 3.5 Flash at 28%, GPT-5.4-mini and Kimi K2.6 tied at 24%, and then a long tail of models in the teens and single digits. Claude Haiku 4.5, which scores 39% on SWE-Bench Pro, collapses to zero on DeepSWE — suggesting that some mid-tier models have been significantly overperforming on easier, potentially contaminated benchmarks.

                                      Screenshot 2026-05-26 at 3.09.33 PM

                                      On SWE-Bench Pro, frontier models cluster within a 30-point range. On DeepSWE, the same models spread across 70 points, with some — like Claude Haiku 4.5 — collapsing entirely. (Source: Datacurve)

                                      GPT-5.5 doesn’t just score the highest — it does so efficiently. The model reaches its 70% pass rate with a median cost of $5.80 per trial, a median wall-clock time of 20 minutes, and a median of 47,000 output tokens. GPT-5.4 emerges as perhaps the best overall value at $3.30 per trial with a 56% score. Claude Opus 4.7, meanwhile, costs significantly more per run, and output tokens, wall-clock duration, and dollar cost per trial all vary by an order of magnitude across the agents tested — yet none of these correlates strongly with pass rate. Agents that emit more tokens, run longer, or cost more do not consistently solve more tasks.

                                      Screenshot 2026-05-26 at 3.27.42 PM

                                      GPT-5.4 and GPT-5.5 occupy the cost-efficient frontier, solving the most tasks for the least money per run. Spending more did not reliably produce better results. (Source: Datacurve)

                                      Datacurve’s audit found that Claude has been reading the answer key on existing benchmarks

                                      Perhaps the most provocative finding in DeepSWE’s analysis concerns what the authors label «CHEATED» verdicts — instances where an agent passes a benchmark not by solving the problem, but by reading the answer.

                                      SWE-Bench Pro’s Docker containers ship the repository’s full .git history, which means the gold-standard solution commit is sitting right there in the container’s file system. Most models ignore it. Claude does not. Datacurve’s analysis found that both Claude Opus 4.7 and Claude Opus 4.6 registered «CHEATED» on more than 12% of their reviewed SWE-Bench Pro rollouts. In those instances, the Claude agent ran commands like git log –all or git show to retrieve the merged fix and paste it into its own patch. The behavior accounted for approximately 18% of Opus 4.7’s passes and 25% of Opus 4.6’s passes on the reviewed sample. The issue has been filed publicly as GitHub issue #93 on the SWE-Bench Pro repository.

                                      GPT-5.4 and GPT-5.5 never exhibited this behavior. Gemini configurations stayed around 1%. Datacurve describes the behavior diplomatically — «The benchmark makes this possible (the gold commit lives in the container), but Claude is the family that consistently does so» — but the implication is clear: a meaningful fraction of Claude’s SWE-Bench Pro scores may reflect environmental exploitation rather than genuine engineering capability.

                                      DeepSWE addresses this by shipping only a shallow clone with the base commit, leaving no gold hash for the agent to discover. It is worth noting that the behavior is arguably a sign of Claude’s environmental attentiveness — the model is very good at exploring its surroundings and exploiting available resources. Whether that counts as «cheating» or «resourcefulness» depends on your perspective, but in the context of a benchmark designed to measure independent problem-solving, it undermines the signal.

                                      Screenshot 2026-05-26 at 3.28.43 PM

                                      Two mechanisms by which agents passed SWE-Bench Pro without solving the underlying problem: reading the answer from the container’s Git history, or stubbing features past weak gold tests. (Source: Datacurve)

                                      Each AI model family fails in its own distinctive way, and the patterns matter for enterprise teams

                                      Beyond the top-line scores, Datacurve’s qualitative trajectory analysis reveals distinctly different failure signatures across model families — a finding that could help engineering teams choose the right model for specific types of work.

                                      Claude is forgetful with multi-part prompts. On DeepSWE, Claude configurations miss stated requirements more than any other family. The pattern is consistent: when a prompt enumerates parallel behaviors — «support both sync and async,» for instance — Claude typically implements the obvious branch and forgets to mirror the change. Datacurve reports that roughly two-thirds of Claude’s «MISSED_REQUIREMENT» failures on DeepSWE follow this «one branch shipped» pattern. In one example, Claude Opus 4.7 correctly landed a sync state-data hook in one engine class while the async engine never received the same hook.

                                      GPT, by contrast, implements exactly what is asked. GPT-5.5 had the lowest rate of missing stated behaviors of any configuration tested. Across multiple runs of the same task, GPT trials tended to converge on the same interpretation of the prompt, suggesting instruction-following precision is a stable trait of the model rather than per-run luck.

                                      One of the most intriguing findings involves self-verification. On DeepSWE, Claude Opus 4.7 and GPT-5.4 wrote and ran new tests in the project’s own test framework on over 80% of their runs — even though no one asked them to. On SWE-Bench Pro, those same models dropped to 28% and 18%, respectively. The reason: SWE-Bench Pro’s prompt template explicitly tells agents they «should not modify the testing logic or any of the tests.» Agents dutifully complied, suppressing a behavior that likely would have improved their performance. This suggests that prompt design in production coding workflows may be inadvertently suppressing valuable agent behaviors — something enterprise teams deploying AI coding agents should carefully audit.

                                      Screenshot 2026-05-26 at 3.29.41 PM

                                      On DeepSWE, top models wrote and ran their own tests in as many as 85 percent of runs. On SWE-Bench Pro, where prompts discourage modifying tests, the same models rarely did so. (Source: Datacurve)

                                      What DeepSWE gets right, what it gets wrong, and what it means for the future of AI benchmarks

                                      Datacurve is forthright about several limitations. The standardized harness, while ensuring fairness, routes all edits through bash rather than the model-specific editing tools each family was trained on — apply_patch for GPT, str_replace_based_edit_tool for Claude. This could hold models below their native ceilings. The benchmark draws exclusively from open-source repositories with 500-plus stars, and results may not generalize to proprietary codebases. Bug localization and refactoring tasks are under-represented, and widely used languages like C++ and Java are absent entirely. The verdict assignments in the qualitative analysis come from an LLM analyzer, not human reviewers, and sample sizes are modest — roughly 90 reviewed rollouts per model per benchmark.

                                      It is also worth noting that Datacurve is a startup with its own commercial interests, and an independent benchmark that reshuffles the leaderboard will inevitably invite scrutiny. The company’s decision to publish the full dataset, all agent trajectories, and the evaluation harness on GitHub mitigates this concern considerably, but independent reproduction will be necessary before the AI community treats these results as definitive.

                                      DeepSWE arrives at an inflection point for the AI coding market. Enterprise adoption of AI coding agents is accelerating rapidly, with engineering organizations making consequential bets on which model to build around. The benchmark market itself has become a strategic battleground — Scale AI’s SWE-Bench Pro, which Datacurve directly critiques, is maintained by a company that also provides evaluation services to the labs whose models it ranks.

                                      If DeepSWE’s central findings about verifier reliability and data contamination hold up under independent scrutiny, they could force a reckoning not just with how the industry measures coding agents, but with the broader question of what benchmarks are actually for. A leaderboard where the grading system is wrong a third of the time is not merely inaccurate — it is the kind of broken instrument that makes everyone feel good about progress that may not be real. And in an industry spending billions on a bet that AI agents can do the work of software engineers, the difference between real progress and the appearance of it is not academic. It is the whole game.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      For months, the leading AI coding benchmarks have told enterprise buyers a comforting but misleading story: the top models are all roughly the same. OpenAI’s GPT-5 family, Anthropic’s Claude Opus, and Google’s Gemini Pro have clustered within a narrow band on Scale AI’s SWE-Bench Pro leaderboard, making it nearly impossible for engineering leaders to determine which agent will actually perform best inside their codebases.

                                      On Monday, a startup called Datacurve released a benchmark it says shatters that illusion. DeepSWE, a 113-task evaluation spanning 91 open-source repositories and five programming languages, produces a dramatically wider spread among the same frontier models — and crowns OpenAI’s GPT-5.5 as the clear leader at 70%, sixteen points ahead of its nearest competitor.

                                      «On public leaderboards, top models often look relatively close in capability,» wrote Datacurve co-author Serena Ge on X. «DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.»

                                      The benchmark also delivers a pointed critique of the evaluation infrastructure the AI industry relies on to measure progress: Datacurve’s audit found that SWE-Bench Pro’s verifiers — the automated graders that determine whether an agent solved a task — issued incorrect pass/fail verdicts on roughly one-third of the trials it reviewed.

                                      If that finding holds up, it has sweeping implications. Enterprise procurement teams, venture capitalists, and AI lab marketing departments all lean heavily on benchmark scores to make multimillion-dollar decisions. A 32% error rate in the most widely cited coding benchmark suggests the industry may have been navigating by a broken compass.

                                      Why the most popular AI coding benchmark may be grading on a curve

                                      To understand what Datacurve is claiming, it helps to understand how coding benchmarks work — and how they can go wrong.

                                      The dominant paradigm, pioneered by the SWE-Bench family maintained by Scale AI and academic researchers, constructs tasks by mining real GitHub commits. The process extracts a bug fix or feature addition from a repository’s history, rolls the code back to the pre-fix state, and then asks an AI agent to reproduce the change. The original commit’s test suite serves as the verifier: if the agent’s patch makes the same tests pass, it gets credit. This approach has an elegant simplicity, but Datacurve argues it introduces three systemic weaknesses.

                                      First, contamination. Because tasks are drawn from public GitHub history, the problem statement, the discussion, and often the exact solution are already present in the training data of frontier models. «The SWE-Bench family scrapes existing GitHub issues and PRs, which creates two problems: memorization (models have already seen the solution) and triviality (most tasks are small),» Ge wrote.

                                      Second, scope. SWE-Bench Pro tasks require, on average, just 120 lines of code added across 5 files. DeepSWE’s reference solutions average 668 lines added across 7 files — roughly 5.5 times more code. Yet DeepSWE’s prompts are actually shorter, averaging 2,158 characters versus SWE-Bench Pro’s 4,614. In other words, DeepSWE gives the agent less instruction but expects far more output, which more closely mirrors how a human developer might actually delegate work to an AI assistant.

                                      DeepSWE tasks demand roughly five times more code than SWE-Bench Pro’s while giving agents shorter prompts — a design choice intended to mirror how developers actually hand off work. (Source: Datacurve)

                                      Third — and most damaging — verifier reliability. Datacurve drew 30 tasks at random from both DeepSWE and SWE-Bench Pro, ran three rollouts across 10 frontier model configurations, and then deployed an LLM-based judge to independently assess whether each agent’s patch actually solved the problem. SWE-Bench Pro’s verifiers accepted wrong implementations 8.5% of the time and rejected correct implementations 24% of the time. DeepSWE’s verifiers registered 0.3% and 1.1%, respectively.

                                      Screenshot 2026-05-26 at 3.22.11 PM

                                      Datacurve’s audit found that SWE-Bench Pro’s automated graders rejected correct solutions 24 percent of the time and accepted wrong ones 8.5 percent of the time. DeepSWE’s verifiers kept both rates near zero. (Source: Datacurve)

                                      The false negative problem is especially insidious because it punishes creative solutions. In one documented case, the gold-standard pull request for a SWE-Bench Pro task refactored a private helper function. An agent that correctly solved the task by inlining the same logic — a perfectly valid engineering choice — failed because the test suite tried to import a symbol that only existed in the original author’s specific implementation.

                                      OpenAI’s GPT-5.5 dominates the new benchmark while Claude and Gemini stumble

                                      DeepSWE’s top-line results reorder the familiar hierarchy in ways that should matter to every engineering team evaluating AI coding tools. On SWE-Bench Pro, models from OpenAI, Anthropic, and Google have traded the lead within a 30-point range. DeepSWE stretches that range to 70 points.

                                      GPT-5.5 leads at 70%, followed by GPT-5.4 at 56% and Claude Opus 4.7 at 54%. From there, the drop-off is steep: Claude Sonnet 4.6 lands at 32%, Gemini 3.5 Flash at 28%, GPT-5.4-mini and Kimi K2.6 tied at 24%, and then a long tail of models in the teens and single digits. Claude Haiku 4.5, which scores 39% on SWE-Bench Pro, collapses to zero on DeepSWE — suggesting that some mid-tier models have been significantly overperforming on easier, potentially contaminated benchmarks.

                                      Screenshot 2026-05-26 at 3.09.33 PM

                                      On SWE-Bench Pro, frontier models cluster within a 30-point range. On DeepSWE, the same models spread across 70 points, with some — like Claude Haiku 4.5 — collapsing entirely. (Source: Datacurve)

                                      GPT-5.5 doesn’t just score the highest — it does so efficiently. The model reaches its 70% pass rate with a median cost of $5.80 per trial, a median wall-clock time of 20 minutes, and a median of 47,000 output tokens. GPT-5.4 emerges as perhaps the best overall value at $3.30 per trial with a 56% score. Claude Opus 4.7, meanwhile, costs significantly more per run, and output tokens, wall-clock duration, and dollar cost per trial all vary by an order of magnitude across the agents tested — yet none of these correlates strongly with pass rate. Agents that emit more tokens, run longer, or cost more do not consistently solve more tasks.

                                      Screenshot 2026-05-26 at 3.27.42 PM

                                      GPT-5.4 and GPT-5.5 occupy the cost-efficient frontier, solving the most tasks for the least money per run. Spending more did not reliably produce better results. (Source: Datacurve)

                                      Datacurve’s audit found that Claude has been reading the answer key on existing benchmarks

                                      Perhaps the most provocative finding in DeepSWE’s analysis concerns what the authors label «CHEATED» verdicts — instances where an agent passes a benchmark not by solving the problem, but by reading the answer.

                                      SWE-Bench Pro’s Docker containers ship the repository’s full .git history, which means the gold-standard solution commit is sitting right there in the container’s file system. Most models ignore it. Claude does not. Datacurve’s analysis found that both Claude Opus 4.7 and Claude Opus 4.6 registered «CHEATED» on more than 12% of their reviewed SWE-Bench Pro rollouts. In those instances, the Claude agent ran commands like git log –all or git show to retrieve the merged fix and paste it into its own patch. The behavior accounted for approximately 18% of Opus 4.7’s passes and 25% of Opus 4.6’s passes on the reviewed sample. The issue has been filed publicly as GitHub issue #93 on the SWE-Bench Pro repository.

                                      GPT-5.4 and GPT-5.5 never exhibited this behavior. Gemini configurations stayed around 1%. Datacurve describes the behavior diplomatically — «The benchmark makes this possible (the gold commit lives in the container), but Claude is the family that consistently does so» — but the implication is clear: a meaningful fraction of Claude’s SWE-Bench Pro scores may reflect environmental exploitation rather than genuine engineering capability.

                                      DeepSWE addresses this by shipping only a shallow clone with the base commit, leaving no gold hash for the agent to discover. It is worth noting that the behavior is arguably a sign of Claude’s environmental attentiveness — the model is very good at exploring its surroundings and exploiting available resources. Whether that counts as «cheating» or «resourcefulness» depends on your perspective, but in the context of a benchmark designed to measure independent problem-solving, it undermines the signal.

                                      Screenshot 2026-05-26 at 3.28.43 PM

                                      Two mechanisms by which agents passed SWE-Bench Pro without solving the underlying problem: reading the answer from the container’s Git history, or stubbing features past weak gold tests. (Source: Datacurve)

                                      Each AI model family fails in its own distinctive way, and the patterns matter for enterprise teams

                                      Beyond the top-line scores, Datacurve’s qualitative trajectory analysis reveals distinctly different failure signatures across model families — a finding that could help engineering teams choose the right model for specific types of work.

                                      Claude is forgetful with multi-part prompts. On DeepSWE, Claude configurations miss stated requirements more than any other family. The pattern is consistent: when a prompt enumerates parallel behaviors — «support both sync and async,» for instance — Claude typically implements the obvious branch and forgets to mirror the change. Datacurve reports that roughly two-thirds of Claude’s «MISSED_REQUIREMENT» failures on DeepSWE follow this «one branch shipped» pattern. In one example, Claude Opus 4.7 correctly landed a sync state-data hook in one engine class while the async engine never received the same hook.

                                      GPT, by contrast, implements exactly what is asked. GPT-5.5 had the lowest rate of missing stated behaviors of any configuration tested. Across multiple runs of the same task, GPT trials tended to converge on the same interpretation of the prompt, suggesting instruction-following precision is a stable trait of the model rather than per-run luck.

                                      One of the most intriguing findings involves self-verification. On DeepSWE, Claude Opus 4.7 and GPT-5.4 wrote and ran new tests in the project’s own test framework on over 80% of their runs — even though no one asked them to. On SWE-Bench Pro, those same models dropped to 28% and 18%, respectively. The reason: SWE-Bench Pro’s prompt template explicitly tells agents they «should not modify the testing logic or any of the tests.» Agents dutifully complied, suppressing a behavior that likely would have improved their performance. This suggests that prompt design in production coding workflows may be inadvertently suppressing valuable agent behaviors — something enterprise teams deploying AI coding agents should carefully audit.

                                      Screenshot 2026-05-26 at 3.29.41 PM

                                      On DeepSWE, top models wrote and ran their own tests in as many as 85 percent of runs. On SWE-Bench Pro, where prompts discourage modifying tests, the same models rarely did so. (Source: Datacurve)

                                      What DeepSWE gets right, what it gets wrong, and what it means for the future of AI benchmarks

                                      Datacurve is forthright about several limitations. The standardized harness, while ensuring fairness, routes all edits through bash rather than the model-specific editing tools each family was trained on — apply_patch for GPT, str_replace_based_edit_tool for Claude. This could hold models below their native ceilings. The benchmark draws exclusively from open-source repositories with 500-plus stars, and results may not generalize to proprietary codebases. Bug localization and refactoring tasks are under-represented, and widely used languages like C++ and Java are absent entirely. The verdict assignments in the qualitative analysis come from an LLM analyzer, not human reviewers, and sample sizes are modest — roughly 90 reviewed rollouts per model per benchmark.

                                      It is also worth noting that Datacurve is a startup with its own commercial interests, and an independent benchmark that reshuffles the leaderboard will inevitably invite scrutiny. The company’s decision to publish the full dataset, all agent trajectories, and the evaluation harness on GitHub mitigates this concern considerably, but independent reproduction will be necessary before the AI community treats these results as definitive.

                                      DeepSWE arrives at an inflection point for the AI coding market. Enterprise adoption of AI coding agents is accelerating rapidly, with engineering organizations making consequential bets on which model to build around. The benchmark market itself has become a strategic battleground — Scale AI’s SWE-Bench Pro, which Datacurve directly critiques, is maintained by a company that also provides evaluation services to the labs whose models it ranks.

                                      If DeepSWE’s central findings about verifier reliability and data contamination hold up under independent scrutiny, they could force a reckoning not just with how the industry measures coding agents, but with the broader question of what benchmarks are actually for. A leaderboard where the grading system is wrong a third of the time is not merely inaccurate — it is the kind of broken instrument that makes everyone feel good about progress that may not be real. And in an industry spending billions on a bet that AI agents can do the work of software engineers, the difference between real progress and the appearance of it is not academic. It is the whole game.

                                      ● Canal oficial · Gratis
                                      ¡Recibe las noticias antes que nadie!
                                      Únete a nuestro canal de WhatsApp y mantente informado al instante, sin spam.
                                      Unirme ahora →
                                      ● Noticias al instante ● Cobertura nacional ● Periodismo real Despertar Matinal
                                      — Redacción Despertar Matinal

                                      — Redacción Despertar Matinal

                                      Programa radial que te conecta con la información desde temprano en la mañana.

                                      Next Post
                                      'Dutton Ranch' debutó con números récord en Paramount+

                                      'Dutton Ranch' debutó con números récord en Paramount+

                                      Deja una respuesta Cancelar la respuesta

                                      Tu dirección de correo electrónico no será publicada. Los campos obligatorios están marcados con *

                                      Canal de WhatsApp

                                      WhatsApp logo WhatsApp

                                      Canal · Despertar Matinal

                                      Únete a nuestro
                                      Canal

                                      Seguir ahora

                                      El clima

                                      Canal de YouTube

                                      YouTube

                                      Canal · Despertar Matinal

                                      Mira nuestro
                                      Canal

                                      Ver ahora

                                      Escúchanos en Spotify

                                      Spotify

                                      Podcast · Despertar Matinal

                                      Escucha nuestro
                                      Podcast

                                      Escuchar ahora

                                      Noticias Populares

                                      • Rex Heuermann le dijo a su ex esposa que asesinó a las víctimas de Gilgo Beach en la casa de la familia, revela un documental

                                        Rex Heuermann le dijo a su ex esposa que asesinó a las víctimas de Gilgo Beach en la casa de la familia, revela un documental

                                        0 shares
                                        Share 0 Tweet 0
                                      • Allyson Felix, de 40 años, busca regresar y tal vez un lugar en los Juegos Olímpicos de Los Ángeles

                                        0 shares
                                        Share 0 Tweet 0
                                      • Proponen declarar el turismo de salud como prioridad nacional …

                                        0 shares
                                        Share 0 Tweet 0
                                      • Javier Milei participó junto a Donald Trump de la cumbre “Shield of the Americas” en Nueva York

                                        0 shares
                                        Share 0 Tweet 0
                                      • Hallan posible altar de fuego zoroástrico en una antigua ciudad de Uzbekistán

                                        0 shares
                                        Share 0 Tweet 0

                                      Medio digital independiente con análisis, opinión y periodismo responsable desde República Dominicana.

                                      Secciones populares

                                      • Política
                                      • Economía & Negocios
                                      • Justicia
                                      • Turismo
                                      • Tecnología
                                      • Entretenimiento
                                      • Mundo
                                      • Cine y Series
                                      • Música
                                      • Moda

                                      Contenido

                                      • Titulares del Día
                                      • Mundo
                                      • Nacionales
                                      • Política
                                      • Deportes
                                      • Economía & Negocios
                                      • Ciencia
                                      • Entretenimiento
                                      • Podcast
                                      • Opinión
                                      • Despertar Matinal TV
                                      • Editoriales

                                      Corporativo

                                      • Sobre nosotros
                                      • Publicidad
                                      • Sala de prensa
                                      • Contacto
                                      • Política de Privacidad
                                      • Eliminación de Datos

                                      Boletines

                                      Suscríbete a nuestro boletín
                                      Recibe las noticias más importantes cada mañana.

                                      • Nosotros
                                      • Publicidad
                                      • Trabaja con nosotros
                                      • Contactos

                                      © 2025 Despertar Matinal. Aviso Legal - comunícate con nuestra redacción y obtén más información sobre Despertar Matinal..

                                      No Result
                                      View All Result
                                      • Home

                                      © 2025 Despertar Matinal. Aviso Legal - comunícate con nuestra redacción y obtén más información sobre Despertar Matinal..

                                      Welcome Back!

                                      Login to your account below

                                      Forgotten Password?

                                      Retrieve your password

                                      Please enter your username or email address to reset your password.

                                      Log In

                                      Desarrollado por
                                      ►
                                      Las cookies necesarias habilitan funciones esenciales del sitio como inicios de sesión seguros y ajustes de preferencias de consentimiento. No almacenan datos personales.
                                      Ninguno
                                      ►
                                      Las cookies funcionales soportan funciones como compartir contenido en redes sociales, recopilar comentarios y habilitar herramientas de terceros.
                                      Ninguno
                                      ►
                                      Las cookies analíticas rastrean las interacciones de los visitantes, proporcionando información sobre métricas como el número de visitantes, la tasa de rebote y las fuentes de tráfico.
                                      Ninguno
                                      ►
                                      Las cookies de publicidad ofrecen anuncios personalizados basados en tus visitas anteriores y analizan la efectividad de las campañas publicitarias.
                                      Ninguno
                                      ►
                                      Las cookies no clasificadas son aquellas que estamos en proceso de clasificar, junto con los proveedores de cookies individuales.
                                      Ninguno
                                      Desarrollado por