• Nosotros
  • Publicidad
  • Trabaja con nosotros
  • Contactos
domingo, septiembre 6, 2026
  • Login
No Result
View All Result
NEWSLETTER
Despertar Matinal
  • Titulares del Día
    • All
    • En Portada
    Inflación en República Dominicana baja a 5.13 %, según Banco Central

    Inflación en República Dominicana baja a 5.13 %, según Banco Central

    Leonel: “El PRM está viniendo a la Fuerza del Pueblo y eso está ocurriendo en todo el país”

    Leonel: “El PRM está viniendo a la Fuerza del Pueblo y eso está ocurriendo en todo el país”

    Presidente Abinader entrega 1,750 títulos de propiedad en Los Alcarrizos que benefician a unas 7,000 personas

    Presidente Abinader entrega 1,750 títulos de propiedad en Los Alcarrizos que benefician a unas 7,000 personas

    Zoraima Cuello afirma ciudadanos se quejan por apagones, delincuencia y alto costo de alimentos

    Zoraima Cuello afirma ciudadanos se quejan por apagones, delincuencia y alto costo de alimentos

    “Me designaron para corregir, no para aferrarme a problemas heredados”, aclara el director del INABIE

    “Me designaron para corregir, no para aferrarme a problemas heredados”, aclara el director del INABIE

    TC devuelve al TSE caso sobre regulación de encuestas y cuestiona remisión del expediente

    TC devuelve al TSE caso sobre regulación de encuestas y cuestiona remisión del expediente

    Presidente Abinader informa negociaciones con Gobierno de Estados Unidos para proteger producción nacional de arroz

    Presidente Abinader informa negociaciones con Gobierno de Estados Unidos para proteger producción nacional de arroz

    Gasolina premium y el gasoil óptimo recibirán reajustes al alza de RD$9.00 cada uno y la gasolina y el gasoil regular de RD$7.00

    Gobierno decide mantener sin variación los precios de los combustibles

    Tras 25 años de espera, presidente Abinader inaugura la carretera de Boca de Chavón

    Tras 25 años de espera, presidente Abinader inaugura la carretera de Boca de Chavón

    Trending Tags

    • Mundo
      • All
      • América Latina
      • Conflictos Internacionales
      • Estados Unidos
      • Europa
      • Geopolítica
      • Haití
      • Medio Oriente
      Vivir gratis en Alemania: buscan voluntarios y ofrecen alojamiento y comida a cambio de trabajar en el campo

      Vivir gratis en Alemania: buscan voluntarios y ofrecen alojamiento y comida a cambio de trabajar en el campo

      Cinco países europeos preparan acuerdos para deportar migrantes fuera de la Unión Europea

      Cinco países europeos preparan acuerdos para deportar migrantes fuera de la Unión Europea

      Tuli Acosta habló de los rumores de romance con Luck Ra: "No me gusta que me digan Tatiana"

      Tuli Acosta habló de los rumores de romance con Luck Ra: «No me gusta que me digan Tatiana»

      Witkoff y Kushner se reunieron con Putin e impulsan las negociaciones de Trump para poner fin a la guerra en Ucrania

      Witkoff y Kushner se reunieron con Putin e impulsan las negociaciones de Trump para poner fin a la guerra en Ucrania

      Por qué presenciamos una Cadena Nacional Histórica

      Por qué presenciamos una Cadena Nacional Histórica

      La petrolera Halliburton confirmó que no participará de ninguna actividad en las Islas Malvinas

      La petrolera Halliburton confirmó que no participará de ninguna actividad en las Islas Malvinas

      Excelente Franco Colapinto: largará 7° en el GP de Italia y se refirió a la pole position conseguida por Pierre Gasly

      Excelente Franco Colapinto: largará 7° en el GP de Italia y se refirió a la pole position conseguida por Pierre Gasly

      La empresa de servicios petroleros más grande del mundo anunció que no participará en ninguna actividad en Malvinas

      La empresa de servicios petroleros más grande del mundo anunció que no participará en ninguna actividad en Malvinas

      INCUCAI: Gracias a Milei el sistema de donación y trasplante se fortalece para dar respuesta a quienes esperan

      INCUCAI: Gracias a Milei el sistema de donación y trasplante se fortalece para dar respuesta a quienes esperan

      Trending Tags

      • Nacionales
        • All
        • Bávaro Punta Cana
        • Educación
        • Gobierno
        • Infraestructura
        • Justicia
        • Obras Públicas
        • Opinión
        • Provincias
        • Seguridad Ciudadana
        • semana santa 2026
        • Sociedad
        • Transporte
        Inflación en República Dominicana baja a 5.13 %, según Banco Central

        Inflación en República Dominicana baja a 5.13 %, según Banco Central

        Milton Ray Guevara advierte riesgos del descalabro en el Poder Judicial para RD

        Milton Ray Guevara advierte riesgos del descalabro en el Poder Judicial para RD

        PN pone en marcha “Ruta Azul” con 84 agentes para reforzar...

        PN pone en marcha “Ruta Azul” con 84 agentes para reforzar…

        Condenan otros siete integraban red de narcotráfico y lavados...

        Condenas de 40 y 30 años de prisión a tres hombres por el…

        Intec celebra Jornada Científica sobre neurociencias dedicada al doctor José Joaquín Puello

        Intec celebra Jornada Científica sobre neurociencias dedicada al doctor José Joaquín Puello

        Leonel: “El PRM está viniendo a la Fuerza del Pueblo y eso está ocurriendo en todo el país”

        Leonel: “El PRM está viniendo a la Fuerza del Pueblo y eso está ocurriendo en todo el país”

        Ministro de Trabajo dice vieja cultura del trabajo infantil ha sido...

        Ministro de Trabajo dice vieja cultura del trabajo infantil ha sido…

        Justicia y Transparencia respalda adhesión de RD al Consenso...

        Justicia y Transparencia respalda adhesión de RD al Consenso…

        Presidente Abinader entrega 1,750 títulos de propiedad en Los Alcarrizos que benefician a unas 7,000 personas

        Presidente Abinader entrega 1,750 títulos de propiedad en Los Alcarrizos que benefician a unas 7,000 personas

        Trending Tags

        • Política
          • All
          • Congreso
          • Opinión Política
          • Partidos Políticos
          • Poder Municipal
          • Transparencia y Corrupción
          Leonel Fernández desarrollará amplia agenda este fin de semana en Azua, San Juan y Elías Piña

          Leonel Fernández desarrollará amplia agenda este fin de semana en Azua, San Juan y Elías Piña

          Fuerza del Pueblo denuncia aumentan las quejas por facturación elevada y persisten los apagones

          Fuerza del Pueblo denuncia aumentan las quejas por facturación elevada y persisten los apagones

          ¡Leonel Fernández llega a San Juan! La Fuerza del Pueblo prepara gran acto de juramentación de nuevos miembros

          ¡Leonel Fernández llega a San Juan! La Fuerza del Pueblo prepara gran acto de juramentación de nuevos miembros

          Dicen proyecto presidencial de Gonzalo Castillo impacta más de 40 territorios el fin de semana

          Dicen proyecto presidencial de Gonzalo Castillo impacta más de 40 territorios el fin de semana

          Antonio Marte juramenta nuevas estructuras fortalecen PPG

          Antonio Marte juramenta nuevas estructuras fortalecen PPG

          Leonel Fernández encabeza asambleas provinciales en despliegue nacional de la Fuerza del Pueblo

          Leonel Fernández encabeza asambleas provinciales en despliegue nacional de la Fuerza del Pueblo

          Rafael Méndez “El Conde” anuncia aspiración a regidor por Fuerza del Pueblo en San Juan de la Maguana

          Rafael Méndez “El Conde” anuncia aspiración a regidor por Fuerza del Pueblo en San Juan de la Maguana

          Vicealcaldesa de Los Alcarrizos abandona el PRM y se juramenta en la Fuerza del Pueblo junto a más de 300 dirigentes

          Vicealcaldesa de Los Alcarrizos abandona el PRM y se juramenta en la Fuerza del Pueblo junto a más de 300 dirigentes

          Leonel afirma PRM no pudo mantener 24 horas de electricidad y la gente está cansada de apagones

          Leonel afirma PRM no pudo mantener 24 horas de electricidad y la gente está cansada de apagones

          Trending Tags

          • Deportes
            • All
            • Atletas Dominicanos
            • Béisbol
            DR Open Kiteboarding Championship reúne atletas de 15 países y reafirma a Cabarete como capital del kitesurf del Caribe

            Cabarete se corona como capital histórica del kitesurf con el DR Open Championship 2026

            El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

            El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

            La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

            La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

            Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

            Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

            Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

            Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

            Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

            Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

            Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

            Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

            30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

            Buffalo recibe a Montreal para abrir la segunda ronda

            Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

            Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

            Trending Tags

            • Economía
              • All
              • Combustibles
              • Energía
              • Indicadores Económicos
              • Sector Energético
              • Turismo
              Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

              Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

              Aventúrate RD 2026

              Aventúrate RD 2026 revela agenda oficial y consolida el turismo de aventura dominicano

              WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

              WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

              Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

              Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

              Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

              Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

              Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

              Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

              El ministerio de Turismo cierra con éxito festival gastronómico “Saborea el Paraíso” en Sánchez, Samaná

              El Ministerio de Turismo celebra un exitoso cierre del festival gastronómico «Saborea el Paraíso» en Sánchez, Samaná

              El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

              El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

              Más allá del comercio: los efectos del arancel estadounidense sobre el turismo dominicano

              Arancel de EE.UU. pone a prueba al turismo dominicano y al silencio oficial del gobierno

              Trending Tags

              • Ciencia
                • All
                • Energía
                • Innovación
                • Investigación Científica
                • Salud y Medicina
                • Tecnología Médica
                Dozens feared trapped in collapsed building in Delhi

                Dozens feared trapped in collapsed building in Delhi

                Volcano eruption triggers flight suspensions at Indonesia's main airport

                Volcano eruption triggers flight suspensions at Indonesia’s main airport

                Russia-Ukraine war: US envoys set for Ukraine talks after Putin meeting

                Russia-Ukraine war: US envoys set for Ukraine talks after Putin meeting

                Families lose 'lifeline' after Haitian caregivers let go as TPS status ends

                Families lose ‘lifeline’ after Haitian caregivers let go as TPS status ends

                Majority Russian-speaking city in Ukraine considers language ban in the arts

                Majority Russian-speaking city in Ukraine considers language ban in the arts

                Flyers might have to face more disruption with hotter skies

                Flyers might have to face more disruption with hotter skies

                Leonel Fernández asegura dirigentes del PRM están pasando a la Fuerza del Pueblo “en todo el país”

                Leonel Fernández asegura dirigentes del PRM están pasando a la Fuerza del Pueblo “en todo el país”

                TV presenter Sarah Khalifa sentenced to death in Egypt drugs case

                TV presenter Sarah Khalifa sentenced to death in Egypt drugs case

                Prince William to attend King Harald's funeral in Norway

                Prince William to attend King Harald’s funeral in Norway

                Trending Tags

                • Tecnología
                  • All
                  • Aplicaciones
                  • Inteligencia Artificial
                  El satélite de rescate se acerca al telescopio condenado de la NASA, incluso si no puede salvarlo

                  El satélite de rescate se acerca al telescopio condenado de la NASA, incluso si no puede salvarlo

                  El sitio web de la Casa Blanca estrena 5 videojuegos arcade retro que promueven la agenda de Trump

                  El sitio web de la Casa Blanca estrena 5 videojuegos arcade retro que promueven la agenda de Trump

                  Las cámaras de vigilancia de multitudes se convierten en el objetivo de la campaña de mitad de mandato mientras los votantes se resisten al poder de las empresas de tecnología

                  Las cámaras de vigilancia de multitudes se convierten en el objetivo de la campaña de mitad de mandato mientras los votantes se resisten al poder de las empresas de tecnología

                  'Welcome to the AGI era': OpenAI launches GPT-6 Astra

                  ‘Welcome to the AGI era’: OpenAI launches GPT-6 Astra

                  Los piratas informáticos vinculados a China abrieron puertas traseras a las computadoras portátiles de los ejecutivos a través de USB, explotando una solución que las empresas tenían pero que no estaban usando.

                  Los piratas informáticos vinculados a China abrieron puertas traseras a las computadoras portátiles de los ejecutivos a través de USB, explotando una solución que las empresas tenían pero que no estaban usando.

                  30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                  Científicos que estudian la leche de cucaracha y sonarse la nariz ganan premios Ig Nobel por ciencia peculiar

                  Meta dice que Muse Spark 1.3 tiene un rendimiento de vanguardia, pero sus mejores resultados provienen de un modelo que los desarrolladores aún no pueden utilizar ampliamente.

                  Meta dice que Muse Spark 1.3 tiene un rendimiento de vanguardia, pero sus mejores resultados provienen de un modelo que los desarrolladores aún no pueden utilizar ampliamente.

                  Un par de naves espaciales se acercan a Mercurio después de un viaje de casi una década

                  Un par de naves espaciales se acercan a Mercurio después de un viaje de casi una década

                  Microsoft AI’s MAI-Transcribe-2 undercuts OpenAI, Google and ElevenLabs on price and speed

                  Microsoft AI’s MAI-Transcribe-2 undercuts OpenAI, Google and ElevenLabs on price and speed

                  Trending Tags

                  • Entretenimiento
                    • All
                    • Cine y Series
                    • Cultura Digital
                    • Cultura Popular
                    • Gastronomía
                    • Música
                    Celine Dion está de regreso en París, pero su primera canción sigue siendo "un gran secreto"

                    Celine Dion está de regreso en París, pero su primera canción sigue siendo «un gran secreto»

                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                    La ‘Odisea’ de Emily Wilson se convirtió en un punto de inflamación cultural. Ahora ella está retraduciendo todo.

                    Editor, editor y reportero de Stars and Stripes demandan al Pentágono para impugnar sus despidos

                    Editor, editor y reportero de Stars and Stripes demandan al Pentágono para impugnar sus despidos

                    Muere Peter Cullen, el prolífico actor de doblaje que le dio a Optimus Prime su autoritario barítono

                    Muere Peter Cullen, el prolífico actor de doblaje que le dio a Optimus Prime su autoritario barítono

                    Juez pregunta por qué el Kennedy Center se está moviendo tan rápido para devolver el nombre de Trump al edificio

                    Juez pregunta por qué el Kennedy Center se está moviendo tan rápido para devolver el nombre de Trump al edificio

                    En el conflictivo norte de Nigeria, una animada vida nocturna convive con una policía moral e inseguridad.

                    En el conflictivo norte de Nigeria, una animada vida nocturna convive con una policía moral e inseguridad.

                    Un teatro reinventa la Odisea de Homero a través de la agonía de la guerra de Ucrania

                    Un teatro reinventa la Odisea de Homero a través de la agonía de la guerra de Ucrania

                    El rapero Yung Filly regresará a Gran Bretaña antes del juicio por violación en Australia el próximo año

                    El rapero Yung Filly regresará a Gran Bretaña antes del juicio por violación en Australia el próximo año

                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                    Los británicos tienen la oportunidad de leer las memorias de Jason Arday en las librerías del Reino Unido

                    Trending Tags

                    • Titulares del Día
                      • All
                      • En Portada
                      Inflación en República Dominicana baja a 5.13 %, según Banco Central

                      Inflación en República Dominicana baja a 5.13 %, según Banco Central

                      Leonel: “El PRM está viniendo a la Fuerza del Pueblo y eso está ocurriendo en todo el país”

                      Leonel: “El PRM está viniendo a la Fuerza del Pueblo y eso está ocurriendo en todo el país”

                      Presidente Abinader entrega 1,750 títulos de propiedad en Los Alcarrizos que benefician a unas 7,000 personas

                      Presidente Abinader entrega 1,750 títulos de propiedad en Los Alcarrizos que benefician a unas 7,000 personas

                      Zoraima Cuello afirma ciudadanos se quejan por apagones, delincuencia y alto costo de alimentos

                      Zoraima Cuello afirma ciudadanos se quejan por apagones, delincuencia y alto costo de alimentos

                      “Me designaron para corregir, no para aferrarme a problemas heredados”, aclara el director del INABIE

                      “Me designaron para corregir, no para aferrarme a problemas heredados”, aclara el director del INABIE

                      TC devuelve al TSE caso sobre regulación de encuestas y cuestiona remisión del expediente

                      TC devuelve al TSE caso sobre regulación de encuestas y cuestiona remisión del expediente

                      Presidente Abinader informa negociaciones con Gobierno de Estados Unidos para proteger producción nacional de arroz

                      Presidente Abinader informa negociaciones con Gobierno de Estados Unidos para proteger producción nacional de arroz

                      Gasolina premium y el gasoil óptimo recibirán reajustes al alza de RD$9.00 cada uno y la gasolina y el gasoil regular de RD$7.00

                      Gobierno decide mantener sin variación los precios de los combustibles

                      Tras 25 años de espera, presidente Abinader inaugura la carretera de Boca de Chavón

                      Tras 25 años de espera, presidente Abinader inaugura la carretera de Boca de Chavón

                      Trending Tags

                      • Mundo
                        • All
                        • América Latina
                        • Conflictos Internacionales
                        • Estados Unidos
                        • Europa
                        • Geopolítica
                        • Haití
                        • Medio Oriente
                        Vivir gratis en Alemania: buscan voluntarios y ofrecen alojamiento y comida a cambio de trabajar en el campo

                        Vivir gratis en Alemania: buscan voluntarios y ofrecen alojamiento y comida a cambio de trabajar en el campo

                        Cinco países europeos preparan acuerdos para deportar migrantes fuera de la Unión Europea

                        Cinco países europeos preparan acuerdos para deportar migrantes fuera de la Unión Europea

                        Tuli Acosta habló de los rumores de romance con Luck Ra: "No me gusta que me digan Tatiana"

                        Tuli Acosta habló de los rumores de romance con Luck Ra: «No me gusta que me digan Tatiana»

                        Witkoff y Kushner se reunieron con Putin e impulsan las negociaciones de Trump para poner fin a la guerra en Ucrania

                        Witkoff y Kushner se reunieron con Putin e impulsan las negociaciones de Trump para poner fin a la guerra en Ucrania

                        Por qué presenciamos una Cadena Nacional Histórica

                        Por qué presenciamos una Cadena Nacional Histórica

                        La petrolera Halliburton confirmó que no participará de ninguna actividad en las Islas Malvinas

                        La petrolera Halliburton confirmó que no participará de ninguna actividad en las Islas Malvinas

                        Excelente Franco Colapinto: largará 7° en el GP de Italia y se refirió a la pole position conseguida por Pierre Gasly

                        Excelente Franco Colapinto: largará 7° en el GP de Italia y se refirió a la pole position conseguida por Pierre Gasly

                        La empresa de servicios petroleros más grande del mundo anunció que no participará en ninguna actividad en Malvinas

                        La empresa de servicios petroleros más grande del mundo anunció que no participará en ninguna actividad en Malvinas

                        INCUCAI: Gracias a Milei el sistema de donación y trasplante se fortalece para dar respuesta a quienes esperan

                        INCUCAI: Gracias a Milei el sistema de donación y trasplante se fortalece para dar respuesta a quienes esperan

                        Trending Tags

                        • Nacionales
                          • All
                          • Bávaro Punta Cana
                          • Educación
                          • Gobierno
                          • Infraestructura
                          • Justicia
                          • Obras Públicas
                          • Opinión
                          • Provincias
                          • Seguridad Ciudadana
                          • semana santa 2026
                          • Sociedad
                          • Transporte
                          Inflación en República Dominicana baja a 5.13 %, según Banco Central

                          Inflación en República Dominicana baja a 5.13 %, según Banco Central

                          Milton Ray Guevara advierte riesgos del descalabro en el Poder Judicial para RD

                          Milton Ray Guevara advierte riesgos del descalabro en el Poder Judicial para RD

                          PN pone en marcha “Ruta Azul” con 84 agentes para reforzar...

                          PN pone en marcha “Ruta Azul” con 84 agentes para reforzar…

                          Condenan otros siete integraban red de narcotráfico y lavados...

                          Condenas de 40 y 30 años de prisión a tres hombres por el…

                          Intec celebra Jornada Científica sobre neurociencias dedicada al doctor José Joaquín Puello

                          Intec celebra Jornada Científica sobre neurociencias dedicada al doctor José Joaquín Puello

                          Leonel: “El PRM está viniendo a la Fuerza del Pueblo y eso está ocurriendo en todo el país”

                          Leonel: “El PRM está viniendo a la Fuerza del Pueblo y eso está ocurriendo en todo el país”

                          Ministro de Trabajo dice vieja cultura del trabajo infantil ha sido...

                          Ministro de Trabajo dice vieja cultura del trabajo infantil ha sido…

                          Justicia y Transparencia respalda adhesión de RD al Consenso...

                          Justicia y Transparencia respalda adhesión de RD al Consenso…

                          Presidente Abinader entrega 1,750 títulos de propiedad en Los Alcarrizos que benefician a unas 7,000 personas

                          Presidente Abinader entrega 1,750 títulos de propiedad en Los Alcarrizos que benefician a unas 7,000 personas

                          Trending Tags

                          • Política
                            • All
                            • Congreso
                            • Opinión Política
                            • Partidos Políticos
                            • Poder Municipal
                            • Transparencia y Corrupción
                            Leonel Fernández desarrollará amplia agenda este fin de semana en Azua, San Juan y Elías Piña

                            Leonel Fernández desarrollará amplia agenda este fin de semana en Azua, San Juan y Elías Piña

                            Fuerza del Pueblo denuncia aumentan las quejas por facturación elevada y persisten los apagones

                            Fuerza del Pueblo denuncia aumentan las quejas por facturación elevada y persisten los apagones

                            ¡Leonel Fernández llega a San Juan! La Fuerza del Pueblo prepara gran acto de juramentación de nuevos miembros

                            ¡Leonel Fernández llega a San Juan! La Fuerza del Pueblo prepara gran acto de juramentación de nuevos miembros

                            Dicen proyecto presidencial de Gonzalo Castillo impacta más de 40 territorios el fin de semana

                            Dicen proyecto presidencial de Gonzalo Castillo impacta más de 40 territorios el fin de semana

                            Antonio Marte juramenta nuevas estructuras fortalecen PPG

                            Antonio Marte juramenta nuevas estructuras fortalecen PPG

                            Leonel Fernández encabeza asambleas provinciales en despliegue nacional de la Fuerza del Pueblo

                            Leonel Fernández encabeza asambleas provinciales en despliegue nacional de la Fuerza del Pueblo

                            Rafael Méndez “El Conde” anuncia aspiración a regidor por Fuerza del Pueblo en San Juan de la Maguana

                            Rafael Méndez “El Conde” anuncia aspiración a regidor por Fuerza del Pueblo en San Juan de la Maguana

                            Vicealcaldesa de Los Alcarrizos abandona el PRM y se juramenta en la Fuerza del Pueblo junto a más de 300 dirigentes

                            Vicealcaldesa de Los Alcarrizos abandona el PRM y se juramenta en la Fuerza del Pueblo junto a más de 300 dirigentes

                            Leonel afirma PRM no pudo mantener 24 horas de electricidad y la gente está cansada de apagones

                            Leonel afirma PRM no pudo mantener 24 horas de electricidad y la gente está cansada de apagones

                            Trending Tags

                            • Deportes
                              • All
                              • Atletas Dominicanos
                              • Béisbol
                              DR Open Kiteboarding Championship reúne atletas de 15 países y reafirma a Cabarete como capital del kitesurf del Caribe

                              Cabarete se corona como capital histórica del kitesurf con el DR Open Championship 2026

                              El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

                              El impulso olímpico del billar recibe un impulso de los dos campeones mundiales consecutivos de China

                              La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

                              La reboteadora líder de todos los tiempos de la WNBA, Tina Charles, se retira del baloncesto

                              Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

                              Sabalenka pide boicot si los jugadores no obtienen una mayor parte de los ingresos del Grand Slam

                              Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

                              Los 76ers tienen un cambio breve y luego una noche larga con una derrota aplastante en el Juego 1

                              Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

                              Ex empleado de Stefon Diggs subirá al estrado por segundo día en el juicio por agresión a un jugador de la NFL

                              Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

                              Kansas City es la sede central de la Copa del Mundo y alberga a Inglaterra, Argentina y Holanda, además de 6 partidos.

                              30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                              Buffalo recibe a Montreal para abrir la segunda ronda

                              Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

                              Judge quiere una nueva tradición del Bronx: “¡Los Yankees ganan!” de Sterling. antes de la canción de Sinatra

                              Trending Tags

                              • Economía
                                • All
                                • Combustibles
                                • Energía
                                • Indicadores Económicos
                                • Sector Energético
                                • Turismo
                                Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

                                Aerodom anuncia nuevas rutas aéreas, pero la pregunta de fondo es quién fiscaliza la concesión

                                Aventúrate RD 2026

                                Aventúrate RD 2026 revela agenda oficial y consolida el turismo de aventura dominicano

                                WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

                                WTTC: Una inversión de más de un billón de dólares en viajes y turismo es una muestra de confianza en el futuro del sector

                                Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

                                Una semana para crear en Samaná: Atelier Yubarta busca conectar arte, naturaleza y turismo en Cayo Levantado Resort

                                Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

                                Meta RD 2036: el plan turístico que el Gobierno aplaude sin fiscalización

                                Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

                                Viva Resorts impulsa el turismo interno en República Dominicana con jornada exclusiva en Bayahibe

                                El ministerio de Turismo cierra con éxito festival gastronómico “Saborea el Paraíso” en Sánchez, Samaná

                                El Ministerio de Turismo celebra un exitoso cierre del festival gastronómico «Saborea el Paraíso» en Sánchez, Samaná

                                El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

                                El Consejo Mundial de Viajes y Turismo (WTTC) informa la incorporación de Piñero como miembro global

                                Más allá del comercio: los efectos del arancel estadounidense sobre el turismo dominicano

                                Arancel de EE.UU. pone a prueba al turismo dominicano y al silencio oficial del gobierno

                                Trending Tags

                                • Ciencia
                                  • All
                                  • Energía
                                  • Innovación
                                  • Investigación Científica
                                  • Salud y Medicina
                                  • Tecnología Médica
                                  Dozens feared trapped in collapsed building in Delhi

                                  Dozens feared trapped in collapsed building in Delhi

                                  Volcano eruption triggers flight suspensions at Indonesia's main airport

                                  Volcano eruption triggers flight suspensions at Indonesia’s main airport

                                  Russia-Ukraine war: US envoys set for Ukraine talks after Putin meeting

                                  Russia-Ukraine war: US envoys set for Ukraine talks after Putin meeting

                                  Families lose 'lifeline' after Haitian caregivers let go as TPS status ends

                                  Families lose ‘lifeline’ after Haitian caregivers let go as TPS status ends

                                  Majority Russian-speaking city in Ukraine considers language ban in the arts

                                  Majority Russian-speaking city in Ukraine considers language ban in the arts

                                  Flyers might have to face more disruption with hotter skies

                                  Flyers might have to face more disruption with hotter skies

                                  Leonel Fernández asegura dirigentes del PRM están pasando a la Fuerza del Pueblo “en todo el país”

                                  Leonel Fernández asegura dirigentes del PRM están pasando a la Fuerza del Pueblo “en todo el país”

                                  TV presenter Sarah Khalifa sentenced to death in Egypt drugs case

                                  TV presenter Sarah Khalifa sentenced to death in Egypt drugs case

                                  Prince William to attend King Harald's funeral in Norway

                                  Prince William to attend King Harald’s funeral in Norway

                                  Trending Tags

                                  • Tecnología
                                    • All
                                    • Aplicaciones
                                    • Inteligencia Artificial
                                    El satélite de rescate se acerca al telescopio condenado de la NASA, incluso si no puede salvarlo

                                    El satélite de rescate se acerca al telescopio condenado de la NASA, incluso si no puede salvarlo

                                    El sitio web de la Casa Blanca estrena 5 videojuegos arcade retro que promueven la agenda de Trump

                                    El sitio web de la Casa Blanca estrena 5 videojuegos arcade retro que promueven la agenda de Trump

                                    Las cámaras de vigilancia de multitudes se convierten en el objetivo de la campaña de mitad de mandato mientras los votantes se resisten al poder de las empresas de tecnología

                                    Las cámaras de vigilancia de multitudes se convierten en el objetivo de la campaña de mitad de mandato mientras los votantes se resisten al poder de las empresas de tecnología

                                    'Welcome to the AGI era': OpenAI launches GPT-6 Astra

                                    ‘Welcome to the AGI era’: OpenAI launches GPT-6 Astra

                                    Los piratas informáticos vinculados a China abrieron puertas traseras a las computadoras portátiles de los ejecutivos a través de USB, explotando una solución que las empresas tenían pero que no estaban usando.

                                    Los piratas informáticos vinculados a China abrieron puertas traseras a las computadoras portátiles de los ejecutivos a través de USB, explotando una solución que las empresas tenían pero que no estaban usando.

                                    30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                    Científicos que estudian la leche de cucaracha y sonarse la nariz ganan premios Ig Nobel por ciencia peculiar

                                    Meta dice que Muse Spark 1.3 tiene un rendimiento de vanguardia, pero sus mejores resultados provienen de un modelo que los desarrolladores aún no pueden utilizar ampliamente.

                                    Meta dice que Muse Spark 1.3 tiene un rendimiento de vanguardia, pero sus mejores resultados provienen de un modelo que los desarrolladores aún no pueden utilizar ampliamente.

                                    Un par de naves espaciales se acercan a Mercurio después de un viaje de casi una década

                                    Un par de naves espaciales se acercan a Mercurio después de un viaje de casi una década

                                    Microsoft AI’s MAI-Transcribe-2 undercuts OpenAI, Google and ElevenLabs on price and speed

                                    Microsoft AI’s MAI-Transcribe-2 undercuts OpenAI, Google and ElevenLabs on price and speed

                                    Trending Tags

                                    • Entretenimiento
                                      • All
                                      • Cine y Series
                                      • Cultura Digital
                                      • Cultura Popular
                                      • Gastronomía
                                      • Música
                                      Celine Dion está de regreso en París, pero su primera canción sigue siendo "un gran secreto"

                                      Celine Dion está de regreso en París, pero su primera canción sigue siendo «un gran secreto»

                                      30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                      La ‘Odisea’ de Emily Wilson se convirtió en un punto de inflamación cultural. Ahora ella está retraduciendo todo.

                                      Editor, editor y reportero de Stars and Stripes demandan al Pentágono para impugnar sus despidos

                                      Editor, editor y reportero de Stars and Stripes demandan al Pentágono para impugnar sus despidos

                                      Muere Peter Cullen, el prolífico actor de doblaje que le dio a Optimus Prime su autoritario barítono

                                      Muere Peter Cullen, el prolífico actor de doblaje que le dio a Optimus Prime su autoritario barítono

                                      Juez pregunta por qué el Kennedy Center se está moviendo tan rápido para devolver el nombre de Trump al edificio

                                      Juez pregunta por qué el Kennedy Center se está moviendo tan rápido para devolver el nombre de Trump al edificio

                                      En el conflictivo norte de Nigeria, una animada vida nocturna convive con una policía moral e inseguridad.

                                      En el conflictivo norte de Nigeria, una animada vida nocturna convive con una policía moral e inseguridad.

                                      Un teatro reinventa la Odisea de Homero a través de la agonía de la guerra de Ucrania

                                      Un teatro reinventa la Odisea de Homero a través de la agonía de la guerra de Ucrania

                                      El rapero Yung Filly regresará a Gran Bretaña antes del juicio por violación en Australia el próximo año

                                      El rapero Yung Filly regresará a Gran Bretaña antes del juicio por violación en Australia el próximo año

                                      30 pasajeros son evacuados después de que un crucero encallara en un arrecife en Fiji

                                      Los británicos tienen la oportunidad de leer las memorias de Jason Arday en las librerías del Reino Unido

                                      Trending Tags

                                      No Result
                                      View All Result
                                      Despertar Matinal
                                      No Result
                                      View All Result

                                      AI IQ is here: a new site scores frontier AI models on the human IQ scale. The results are already dividing tech.

                                      by — Redacción Despertar Matinal
                                      13 de mayo de 2026
                                      in Tecnología
                                      0
                                      AI IQ is here: a new site scores frontier AI models on the human IQ scale. The results are already dividing tech.
                                      0
                                      SHARES
                                      19
                                      VIEWS
                                      Share on FacebookShare on Twitter

                                      For decades, the IQ test has been one of the most familiar — and most contested — yardsticks for human intelligence. Now, a startup project called AI IQ is applying the same metaphor to artificial intelligence, assigning estimated intelligence quotients to more than 50 of the world’s most powerful language models and plotting them on a standard bell curve.

                                      The result is a set of interactive visualizations at aiiq.org that have ricocheted across social media in the past week, drawing praise from enterprise technologists who say the charts make an impossibly complex market legible — and sharp criticism from researchers and commentators who warn the entire framework is misleading.

                                      «This is super useful,» wrote Thibaut Mélen, a technology commentator, on X. «Much easier to understand model progress when it’s mapped like this instead of another giant leaderboard table.»

                                      Brian Vellmure, a business strategist, offered a similar endorsement: «This is helpful. Anecdotally tracks with personal experience.»

                                      But the backlash arrived just as quickly. «It’s nonsense. AI is far too jagged. The map is not the territory,» posted AI Deeply, an artificial intelligence commentary account, crystallizing a worry shared by many researchers: that reducing a language model’s sprawling, uneven capabilities to a single number creates a dangerous illusion of precision.

                                      More than 50 AI language models, plotted on a standard IQ bell curve by the site AI IQ. The most capable models crowd the right tail of the distribution. (Credit: AI IQ)

                                      Twelve benchmarks, four dimensions, and one controversial number: how AI IQ actually works

                                      AI IQ was created by Ryan Shea, an engineer, entrepreneur, and angel investor best known as a co-founder of the blockchain platform Stacks. Shea also co-founded Voterbase and has invested in the early stages of several unicorns, including OpenSea, Lattice, Anchorage, and Mercury. He holds a Bachelor of Science in Mechanical Engineering from Princeton University.

                                      The site’s methodology rests on a deceptively simple formula. AI IQ groups 12 benchmarks into four reasoning dimensions: abstract, mathematical, programmatic, and academic. The composite IQ is a straight average of those four dimension scores: IQ = ¼ (IQ_Abstract + IQ_Math + IQ_Prog + IQ_Acad).

                                      The abstract reasoning dimension draws from ARC-AGI-1 and ARC-AGI-2, the notoriously difficult pattern-recognition benchmarks designed to test general fluid intelligence. Mathematical reasoning includes FrontierMath (Tiers 1–3 and Tier 4), AIME, and ProofBench. Programmatic reasoning uses Terminal-Bench 2.0, SWE-Bench Verified, and SciCode. Academic reasoning pulls from Humanity’s Last Exam, CritPt, and GPQA Diamond.

                                      Each raw benchmark score gets mapped to an implied IQ through what the site describes as «hand-calibrated difficulty curves.» Crucially, the methodology compresses ceilings for benchmarks considered easier or more susceptible to data contamination, preventing them from inflating scores above 100. Harder, less gameable benchmarks retain higher ceilings. The system also handles missing data conservatively: models need scores on at least two of the four dimensions to receive a derived IQ, and when benchmarks are absent, the pipeline deliberately pulls scores down rather than up. The site states that «every derived IQ averages all four dimensions, so missing coverage cannot make a model look better by omission.»

                                      OpenAI leads the bell curve, but the gap between the top AI models has never been smaller

                                      As of mid-May 2026, the AI IQ charts tell a story of rapid convergence at the top of the frontier — and widening diversity in the tiers below.

                                      According to the Frontier IQ Over Time chart, GPT-5.5 from OpenAI currently sits at the peak of the bell curve, with an estimated IQ near 136 — the highest of any model tracked. It is closely followed by GPT-5.4 (approximately 131), Opus 4.7 from Anthropic (approximately 132), and Opus 4.6 (approximately 129). Google’s Gemini 3.1 Pro lands near 131, making the top cluster extraordinarily tight.

                                      That compression is not unique to AI IQ’s framework. Visual Capitalist, drawing from a separate Mensa-based ranking by TrackingAI, recently observed the same dynamic, noting that «the biggest takeaway is how compressed the top of the leaderboard has become.» On that scale, Grok-4.20 Expert Mode and GPT 5.4 Pro tied at 145, with Gemini 3.1 Pro at 141.

                                      Below the frontier cluster, the AI IQ charts show a crowded midfield. Models from Chinese labs — Kimi K2.6, GLM-5, DeepSeek-V3.2, Qwen3.6, MiniMax-M2.7 — bunch between roughly 112 and 118, making the cost-performance tier increasingly competitive for enterprise buyers who don’t need the absolute best model for every task. One X user, ovsky, noted that the data «confirms experience with sonnet 4.6 being an absolute workhorse as opposed to opus 4.5» — pointing to the way the charts can validate practitioner intuitions that headline rankings often miss.

                                      aiiq-frontier-iq-over-time-2026-05-13

                                      The trajectory of frontier AI models from October 2023 to mid-2026, as tracked by AI IQ. Provider-colored step-lines connect each lab’s flagship releases, showing roughly 60 points of estimated IQ improvement in 30 months. (Credit: AI IQ)

                                      Why emotional intelligence scores are becoming the new battleground in AI model rankings

                                      What distinguishes AI IQ from most other benchmarking efforts is its inclusion of an «EQ» — emotional intelligence — score. The site maps each model’s EQ-Bench 3 Elo score and Arena Elo score to an estimated EQ using calibrated piecewise-linear scales, then takes a 50/50 weighted composite of the two.

                                      The EQ scores produce a meaningfully different ranking than IQ alone. On the IQ vs. EQ scatter plot, Anthropic’s Opus 4.7 leads on EQ with a score near 132, pushing it into the upper-right quadrant — the most desirable position, signaling both high cognitive and high emotional intelligence. OpenAI’s GPT-5.5 and GPT-5.4 cluster in the high-IQ zone but lag slightly on EQ. Google’s Gemini 3.1 Pro sits in a strong middle position on both axes.

                                      One notable methodological choice has drawn attention: EQ-Bench 3 is judged by Claude, an Anthropic model, which the site acknowledges «creates potential scoring bias in favor of Anthropic models.» To correct for this, AI IQ subtracts a 200-point Elo penalty from the EQ-Bench component for all Anthropic models before mapping to implied EQ. The Arena component is unaffected since it uses human judges. That self-correction is unusual in the benchmarking world, and it suggests Shea is aware of the methodological minefield he has entered. Still, the EQ dimension captures something IQ alone cannot: the growing importance of conversational quality, collaboration, and trust in models deployed for user-facing work.

                                      aiiq-iq-vs-eq-2026-05-13

                                      Plotting IQ against EQ reveals that the smartest models aren’t always the most emotionally intelligent. Anthropic’s Opus 4.7 dominates the upper-right quadrant. (Credit: AI IQ)

                                      The AI cost-performance chart that enterprise buyers actually need to see

                                      Perhaps the most practically useful chart on the site is not the bell curve but the IQ vs. Effective Cost scatter plot. It maps each model’s estimated IQ against an «effective cost» metric — defined as the token cost for a task using 2 million input tokens and 1 million output tokens, multiplied by a usage efficiency factor.

                                      The chart reveals a familiar pattern in enterprise technology: the best models are not always the best value. GPT-5.5 and Opus 4.7 sit in the upper-left corner — high IQ, high cost, with effective per-task costs north of $30 and $50 respectively. Meanwhile, models like GPT-5.4-mini, DeepSeek-V3.2, and MiniMax-M2.7 occupy a sweet spot in the middle: respectable IQ scores between 112 and 120, at effective costs ranging from roughly $1 to $5 per task. At the cheapest extreme, GPT-oss-20b (an open-source OpenAI model) appears near $0.20 effective cost with an IQ around 107 — potentially the most economical option for bulk classification or extraction workloads.

                                      The site also offers a 3D visualization mapping IQ, EQ, and effective cost simultaneously. A dashed line running through the cube points toward the ideal: higher IQ, higher EQ, and lower cost. Models near the «green end» of that axis are stronger all-around deals; those near the «red end» sacrifice capability, cost efficiency, or both. For CIOs staring at API invoices, the implication is clear: the intelligence gap between a $50 model and a $3 model has narrowed enough that routing — using expensive models for hard problems and cheap ones for everything else — is no longer optional. It is the dominant architecture for serious AI deployments.

                                      Critics say AI’s «jagged» capabilities make a single IQ score dangerously misleading

                                      The loudest objection to AI IQ is philosophical, and it cuts deep. Critics argue that collapsing a model’s uneven capabilities into a single score obscures more than it reveals.

                                      «IQ as a proxy is fading — we’re seeing reasoning density spikes that don’t map to g-factor,» posted Zaya, a technology commentator, on X. «GPT-5.5 already hit saturation on MMLU-Pro, but still fails ClockBench 50% of the time.»

                                      That observation touches on what AI researchers call the «jaggedness» problem: large language models often exhibit wildly uneven capabilities, excelling at graduate-level physics while failing at tasks a child could do. A composite score can paper over those gaps.

                                      Pressureangle, another X user, posted a more granular critique, calling out «complete lack of transparency» and arguing the site never fully discloses how its calibration curves were created or validated. In fairness, AI IQ does list its 12 benchmarks and shows the shape of each calibration curve in its methodology modal. But the raw data and precise mathematical transformations are not published as open datasets — a gap that matters to researchers accustomed to fully reproducible methods.

                                      Others questioned the premise itself. «As useless as human IQ testing,» wrote haashim on X. Shubham Sharma, an AI and technology writer, offered a constructive alternative: «Why not having the Models take an official (MENSA-Grade) test? Wouldn’t this be the most accurate and most ‘human-comparable’ way to benchmark intelligence?» That approach already exists through TrackingAI, which administers the Mensa Norway IQ test to language models. But Mensa-style tests measure only abstract pattern recognition, while AI IQ attempts a broader composite across coding, mathematics, and academic reasoning. As Visual Capitalist noted, «an IQ-style benchmark captures only one slice of capability.» Each approach has tradeoffs — and neither has won the argument yet.

                                      The real race isn’t for the highest score — it’s for the smartest model stack

                                      For all the debate about methodology, the most important signal in AI IQ’s data may not be any single model’s score. It is the shape of the market the charts reveal.

                                      There are now more than 50 frontier-class models available through APIs, from at least 14 major providers spanning the United States, China, and Europe. Each provider publishes its own benchmarks, often cherry-picked to showcase strengths. The result is a Tower of Babel where no two companies measure the same thing in the same way. Academic research has highlighted that «most benchmarks introduce bias by focusing on a particular type of domain,» and the Frontier IQ Over Time chart on AI IQ shows just how fast the targets are moving: in October 2023, GPT-4-turbo sat near an estimated IQ of 75. By early 2026, the top models were brushing 135 — roughly 60 points of improvement in 30 months.

                                      That pace raises a fundamental question about whether any scoring system can keep up. The site compresses ceilings for saturated benchmarks, but as models continue to max out even the hardest tests — ARC-AGI-2, FrontierMath Tier 4, Humanity’s Last Exam — the framework will face the same ceiling effects that have plagued every AI evaluation before it. Connor Forsyth pointed to this dynamic on X: «ARC AGI 3 disagrees,» he wrote, referencing a next-generation benchmark that may already be undermining current scores.

                                      AI IQ is not perfect. Its methodology is partially opaque. Its IQ metaphor can mislead. And its creator acknowledges known biases while likely missing others. But the alternative — wading through dozens of provider-specific benchmark tables, each using different test suites and scoring conventions — is worse. The site offers enterprise buyers something genuinely scarce: a single framework for comparing models across providers, dimensions, and price points, updated regularly, with enough nuance to show that the right answer to «which model is best?» is almost always «it depends on the task.»

                                      As Debdoot Ghosh mused on X after viewing the charts: «Now a human’s role is just to orchestrate?«

                                      Maybe. But if the AI IQ data shows anything clearly, it is that orchestration — knowing which model to deploy, when, and at what price — has become its own form of intelligence. And for that, there is no benchmark yet.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      For decades, the IQ test has been one of the most familiar — and most contested — yardsticks for human intelligence. Now, a startup project called AI IQ is applying the same metaphor to artificial intelligence, assigning estimated intelligence quotients to more than 50 of the world’s most powerful language models and plotting them on a standard bell curve.

                                      The result is a set of interactive visualizations at aiiq.org that have ricocheted across social media in the past week, drawing praise from enterprise technologists who say the charts make an impossibly complex market legible — and sharp criticism from researchers and commentators who warn the entire framework is misleading.

                                      «This is super useful,» wrote Thibaut Mélen, a technology commentator, on X. «Much easier to understand model progress when it’s mapped like this instead of another giant leaderboard table.»

                                      Brian Vellmure, a business strategist, offered a similar endorsement: «This is helpful. Anecdotally tracks with personal experience.»

                                      But the backlash arrived just as quickly. «It’s nonsense. AI is far too jagged. The map is not the territory,» posted AI Deeply, an artificial intelligence commentary account, crystallizing a worry shared by many researchers: that reducing a language model’s sprawling, uneven capabilities to a single number creates a dangerous illusion of precision.

                                      More than 50 AI language models, plotted on a standard IQ bell curve by the site AI IQ. The most capable models crowd the right tail of the distribution. (Credit: AI IQ)

                                      Twelve benchmarks, four dimensions, and one controversial number: how AI IQ actually works

                                      AI IQ was created by Ryan Shea, an engineer, entrepreneur, and angel investor best known as a co-founder of the blockchain platform Stacks. Shea also co-founded Voterbase and has invested in the early stages of several unicorns, including OpenSea, Lattice, Anchorage, and Mercury. He holds a Bachelor of Science in Mechanical Engineering from Princeton University.

                                      The site’s methodology rests on a deceptively simple formula. AI IQ groups 12 benchmarks into four reasoning dimensions: abstract, mathematical, programmatic, and academic. The composite IQ is a straight average of those four dimension scores: IQ = ¼ (IQ_Abstract + IQ_Math + IQ_Prog + IQ_Acad).

                                      The abstract reasoning dimension draws from ARC-AGI-1 and ARC-AGI-2, the notoriously difficult pattern-recognition benchmarks designed to test general fluid intelligence. Mathematical reasoning includes FrontierMath (Tiers 1–3 and Tier 4), AIME, and ProofBench. Programmatic reasoning uses Terminal-Bench 2.0, SWE-Bench Verified, and SciCode. Academic reasoning pulls from Humanity’s Last Exam, CritPt, and GPQA Diamond.

                                      Each raw benchmark score gets mapped to an implied IQ through what the site describes as «hand-calibrated difficulty curves.» Crucially, the methodology compresses ceilings for benchmarks considered easier or more susceptible to data contamination, preventing them from inflating scores above 100. Harder, less gameable benchmarks retain higher ceilings. The system also handles missing data conservatively: models need scores on at least two of the four dimensions to receive a derived IQ, and when benchmarks are absent, the pipeline deliberately pulls scores down rather than up. The site states that «every derived IQ averages all four dimensions, so missing coverage cannot make a model look better by omission.»

                                      OpenAI leads the bell curve, but the gap between the top AI models has never been smaller

                                      As of mid-May 2026, the AI IQ charts tell a story of rapid convergence at the top of the frontier — and widening diversity in the tiers below.

                                      According to the Frontier IQ Over Time chart, GPT-5.5 from OpenAI currently sits at the peak of the bell curve, with an estimated IQ near 136 — the highest of any model tracked. It is closely followed by GPT-5.4 (approximately 131), Opus 4.7 from Anthropic (approximately 132), and Opus 4.6 (approximately 129). Google’s Gemini 3.1 Pro lands near 131, making the top cluster extraordinarily tight.

                                      That compression is not unique to AI IQ’s framework. Visual Capitalist, drawing from a separate Mensa-based ranking by TrackingAI, recently observed the same dynamic, noting that «the biggest takeaway is how compressed the top of the leaderboard has become.» On that scale, Grok-4.20 Expert Mode and GPT 5.4 Pro tied at 145, with Gemini 3.1 Pro at 141.

                                      Below the frontier cluster, the AI IQ charts show a crowded midfield. Models from Chinese labs — Kimi K2.6, GLM-5, DeepSeek-V3.2, Qwen3.6, MiniMax-M2.7 — bunch between roughly 112 and 118, making the cost-performance tier increasingly competitive for enterprise buyers who don’t need the absolute best model for every task. One X user, ovsky, noted that the data «confirms experience with sonnet 4.6 being an absolute workhorse as opposed to opus 4.5» — pointing to the way the charts can validate practitioner intuitions that headline rankings often miss.

                                      aiiq-frontier-iq-over-time-2026-05-13

                                      The trajectory of frontier AI models from October 2023 to mid-2026, as tracked by AI IQ. Provider-colored step-lines connect each lab’s flagship releases, showing roughly 60 points of estimated IQ improvement in 30 months. (Credit: AI IQ)

                                      Why emotional intelligence scores are becoming the new battleground in AI model rankings

                                      What distinguishes AI IQ from most other benchmarking efforts is its inclusion of an «EQ» — emotional intelligence — score. The site maps each model’s EQ-Bench 3 Elo score and Arena Elo score to an estimated EQ using calibrated piecewise-linear scales, then takes a 50/50 weighted composite of the two.

                                      The EQ scores produce a meaningfully different ranking than IQ alone. On the IQ vs. EQ scatter plot, Anthropic’s Opus 4.7 leads on EQ with a score near 132, pushing it into the upper-right quadrant — the most desirable position, signaling both high cognitive and high emotional intelligence. OpenAI’s GPT-5.5 and GPT-5.4 cluster in the high-IQ zone but lag slightly on EQ. Google’s Gemini 3.1 Pro sits in a strong middle position on both axes.

                                      One notable methodological choice has drawn attention: EQ-Bench 3 is judged by Claude, an Anthropic model, which the site acknowledges «creates potential scoring bias in favor of Anthropic models.» To correct for this, AI IQ subtracts a 200-point Elo penalty from the EQ-Bench component for all Anthropic models before mapping to implied EQ. The Arena component is unaffected since it uses human judges. That self-correction is unusual in the benchmarking world, and it suggests Shea is aware of the methodological minefield he has entered. Still, the EQ dimension captures something IQ alone cannot: the growing importance of conversational quality, collaboration, and trust in models deployed for user-facing work.

                                      aiiq-iq-vs-eq-2026-05-13

                                      Plotting IQ against EQ reveals that the smartest models aren’t always the most emotionally intelligent. Anthropic’s Opus 4.7 dominates the upper-right quadrant. (Credit: AI IQ)

                                      The AI cost-performance chart that enterprise buyers actually need to see

                                      Perhaps the most practically useful chart on the site is not the bell curve but the IQ vs. Effective Cost scatter plot. It maps each model’s estimated IQ against an «effective cost» metric — defined as the token cost for a task using 2 million input tokens and 1 million output tokens, multiplied by a usage efficiency factor.

                                      The chart reveals a familiar pattern in enterprise technology: the best models are not always the best value. GPT-5.5 and Opus 4.7 sit in the upper-left corner — high IQ, high cost, with effective per-task costs north of $30 and $50 respectively. Meanwhile, models like GPT-5.4-mini, DeepSeek-V3.2, and MiniMax-M2.7 occupy a sweet spot in the middle: respectable IQ scores between 112 and 120, at effective costs ranging from roughly $1 to $5 per task. At the cheapest extreme, GPT-oss-20b (an open-source OpenAI model) appears near $0.20 effective cost with an IQ around 107 — potentially the most economical option for bulk classification or extraction workloads.

                                      The site also offers a 3D visualization mapping IQ, EQ, and effective cost simultaneously. A dashed line running through the cube points toward the ideal: higher IQ, higher EQ, and lower cost. Models near the «green end» of that axis are stronger all-around deals; those near the «red end» sacrifice capability, cost efficiency, or both. For CIOs staring at API invoices, the implication is clear: the intelligence gap between a $50 model and a $3 model has narrowed enough that routing — using expensive models for hard problems and cheap ones for everything else — is no longer optional. It is the dominant architecture for serious AI deployments.

                                      Critics say AI’s «jagged» capabilities make a single IQ score dangerously misleading

                                      The loudest objection to AI IQ is philosophical, and it cuts deep. Critics argue that collapsing a model’s uneven capabilities into a single score obscures more than it reveals.

                                      «IQ as a proxy is fading — we’re seeing reasoning density spikes that don’t map to g-factor,» posted Zaya, a technology commentator, on X. «GPT-5.5 already hit saturation on MMLU-Pro, but still fails ClockBench 50% of the time.»

                                      That observation touches on what AI researchers call the «jaggedness» problem: large language models often exhibit wildly uneven capabilities, excelling at graduate-level physics while failing at tasks a child could do. A composite score can paper over those gaps.

                                      Pressureangle, another X user, posted a more granular critique, calling out «complete lack of transparency» and arguing the site never fully discloses how its calibration curves were created or validated. In fairness, AI IQ does list its 12 benchmarks and shows the shape of each calibration curve in its methodology modal. But the raw data and precise mathematical transformations are not published as open datasets — a gap that matters to researchers accustomed to fully reproducible methods.

                                      Others questioned the premise itself. «As useless as human IQ testing,» wrote haashim on X. Shubham Sharma, an AI and technology writer, offered a constructive alternative: «Why not having the Models take an official (MENSA-Grade) test? Wouldn’t this be the most accurate and most ‘human-comparable’ way to benchmark intelligence?» That approach already exists through TrackingAI, which administers the Mensa Norway IQ test to language models. But Mensa-style tests measure only abstract pattern recognition, while AI IQ attempts a broader composite across coding, mathematics, and academic reasoning. As Visual Capitalist noted, «an IQ-style benchmark captures only one slice of capability.» Each approach has tradeoffs — and neither has won the argument yet.

                                      The real race isn’t for the highest score — it’s for the smartest model stack

                                      For all the debate about methodology, the most important signal in AI IQ’s data may not be any single model’s score. It is the shape of the market the charts reveal.

                                      There are now more than 50 frontier-class models available through APIs, from at least 14 major providers spanning the United States, China, and Europe. Each provider publishes its own benchmarks, often cherry-picked to showcase strengths. The result is a Tower of Babel where no two companies measure the same thing in the same way. Academic research has highlighted that «most benchmarks introduce bias by focusing on a particular type of domain,» and the Frontier IQ Over Time chart on AI IQ shows just how fast the targets are moving: in October 2023, GPT-4-turbo sat near an estimated IQ of 75. By early 2026, the top models were brushing 135 — roughly 60 points of improvement in 30 months.

                                      That pace raises a fundamental question about whether any scoring system can keep up. The site compresses ceilings for saturated benchmarks, but as models continue to max out even the hardest tests — ARC-AGI-2, FrontierMath Tier 4, Humanity’s Last Exam — the framework will face the same ceiling effects that have plagued every AI evaluation before it. Connor Forsyth pointed to this dynamic on X: «ARC AGI 3 disagrees,» he wrote, referencing a next-generation benchmark that may already be undermining current scores.

                                      AI IQ is not perfect. Its methodology is partially opaque. Its IQ metaphor can mislead. And its creator acknowledges known biases while likely missing others. But the alternative — wading through dozens of provider-specific benchmark tables, each using different test suites and scoring conventions — is worse. The site offers enterprise buyers something genuinely scarce: a single framework for comparing models across providers, dimensions, and price points, updated regularly, with enough nuance to show that the right answer to «which model is best?» is almost always «it depends on the task.»

                                      As Debdoot Ghosh mused on X after viewing the charts: «Now a human’s role is just to orchestrate?«

                                      Maybe. But if the AI IQ data shows anything clearly, it is that orchestration — knowing which model to deploy, when, and at what price — has become its own form of intelligence. And for that, there is no benchmark yet.

                                      Tours Colombia Todo el año Tours Colombia Todo el año Tours Colombia Todo el año

                                      For decades, the IQ test has been one of the most familiar — and most contested — yardsticks for human intelligence. Now, a startup project called AI IQ is applying the same metaphor to artificial intelligence, assigning estimated intelligence quotients to more than 50 of the world’s most powerful language models and plotting them on a standard bell curve.

                                      The result is a set of interactive visualizations at aiiq.org that have ricocheted across social media in the past week, drawing praise from enterprise technologists who say the charts make an impossibly complex market legible — and sharp criticism from researchers and commentators who warn the entire framework is misleading.

                                      «This is super useful,» wrote Thibaut Mélen, a technology commentator, on X. «Much easier to understand model progress when it’s mapped like this instead of another giant leaderboard table.»

                                      Brian Vellmure, a business strategist, offered a similar endorsement: «This is helpful. Anecdotally tracks with personal experience.»

                                      But the backlash arrived just as quickly. «It’s nonsense. AI is far too jagged. The map is not the territory,» posted AI Deeply, an artificial intelligence commentary account, crystallizing a worry shared by many researchers: that reducing a language model’s sprawling, uneven capabilities to a single number creates a dangerous illusion of precision.

                                      More than 50 AI language models, plotted on a standard IQ bell curve by the site AI IQ. The most capable models crowd the right tail of the distribution. (Credit: AI IQ)

                                      Twelve benchmarks, four dimensions, and one controversial number: how AI IQ actually works

                                      AI IQ was created by Ryan Shea, an engineer, entrepreneur, and angel investor best known as a co-founder of the blockchain platform Stacks. Shea also co-founded Voterbase and has invested in the early stages of several unicorns, including OpenSea, Lattice, Anchorage, and Mercury. He holds a Bachelor of Science in Mechanical Engineering from Princeton University.

                                      The site’s methodology rests on a deceptively simple formula. AI IQ groups 12 benchmarks into four reasoning dimensions: abstract, mathematical, programmatic, and academic. The composite IQ is a straight average of those four dimension scores: IQ = ¼ (IQ_Abstract + IQ_Math + IQ_Prog + IQ_Acad).

                                      The abstract reasoning dimension draws from ARC-AGI-1 and ARC-AGI-2, the notoriously difficult pattern-recognition benchmarks designed to test general fluid intelligence. Mathematical reasoning includes FrontierMath (Tiers 1–3 and Tier 4), AIME, and ProofBench. Programmatic reasoning uses Terminal-Bench 2.0, SWE-Bench Verified, and SciCode. Academic reasoning pulls from Humanity’s Last Exam, CritPt, and GPQA Diamond.

                                      Each raw benchmark score gets mapped to an implied IQ through what the site describes as «hand-calibrated difficulty curves.» Crucially, the methodology compresses ceilings for benchmarks considered easier or more susceptible to data contamination, preventing them from inflating scores above 100. Harder, less gameable benchmarks retain higher ceilings. The system also handles missing data conservatively: models need scores on at least two of the four dimensions to receive a derived IQ, and when benchmarks are absent, the pipeline deliberately pulls scores down rather than up. The site states that «every derived IQ averages all four dimensions, so missing coverage cannot make a model look better by omission.»

                                      OpenAI leads the bell curve, but the gap between the top AI models has never been smaller

                                      As of mid-May 2026, the AI IQ charts tell a story of rapid convergence at the top of the frontier — and widening diversity in the tiers below.

                                      According to the Frontier IQ Over Time chart, GPT-5.5 from OpenAI currently sits at the peak of the bell curve, with an estimated IQ near 136 — the highest of any model tracked. It is closely followed by GPT-5.4 (approximately 131), Opus 4.7 from Anthropic (approximately 132), and Opus 4.6 (approximately 129). Google’s Gemini 3.1 Pro lands near 131, making the top cluster extraordinarily tight.

                                      That compression is not unique to AI IQ’s framework. Visual Capitalist, drawing from a separate Mensa-based ranking by TrackingAI, recently observed the same dynamic, noting that «the biggest takeaway is how compressed the top of the leaderboard has become.» On that scale, Grok-4.20 Expert Mode and GPT 5.4 Pro tied at 145, with Gemini 3.1 Pro at 141.

                                      Below the frontier cluster, the AI IQ charts show a crowded midfield. Models from Chinese labs — Kimi K2.6, GLM-5, DeepSeek-V3.2, Qwen3.6, MiniMax-M2.7 — bunch between roughly 112 and 118, making the cost-performance tier increasingly competitive for enterprise buyers who don’t need the absolute best model for every task. One X user, ovsky, noted that the data «confirms experience with sonnet 4.6 being an absolute workhorse as opposed to opus 4.5» — pointing to the way the charts can validate practitioner intuitions that headline rankings often miss.

                                      aiiq-frontier-iq-over-time-2026-05-13

                                      The trajectory of frontier AI models from October 2023 to mid-2026, as tracked by AI IQ. Provider-colored step-lines connect each lab’s flagship releases, showing roughly 60 points of estimated IQ improvement in 30 months. (Credit: AI IQ)

                                      Why emotional intelligence scores are becoming the new battleground in AI model rankings

                                      What distinguishes AI IQ from most other benchmarking efforts is its inclusion of an «EQ» — emotional intelligence — score. The site maps each model’s EQ-Bench 3 Elo score and Arena Elo score to an estimated EQ using calibrated piecewise-linear scales, then takes a 50/50 weighted composite of the two.

                                      The EQ scores produce a meaningfully different ranking than IQ alone. On the IQ vs. EQ scatter plot, Anthropic’s Opus 4.7 leads on EQ with a score near 132, pushing it into the upper-right quadrant — the most desirable position, signaling both high cognitive and high emotional intelligence. OpenAI’s GPT-5.5 and GPT-5.4 cluster in the high-IQ zone but lag slightly on EQ. Google’s Gemini 3.1 Pro sits in a strong middle position on both axes.

                                      One notable methodological choice has drawn attention: EQ-Bench 3 is judged by Claude, an Anthropic model, which the site acknowledges «creates potential scoring bias in favor of Anthropic models.» To correct for this, AI IQ subtracts a 200-point Elo penalty from the EQ-Bench component for all Anthropic models before mapping to implied EQ. The Arena component is unaffected since it uses human judges. That self-correction is unusual in the benchmarking world, and it suggests Shea is aware of the methodological minefield he has entered. Still, the EQ dimension captures something IQ alone cannot: the growing importance of conversational quality, collaboration, and trust in models deployed for user-facing work.

                                      aiiq-iq-vs-eq-2026-05-13

                                      Plotting IQ against EQ reveals that the smartest models aren’t always the most emotionally intelligent. Anthropic’s Opus 4.7 dominates the upper-right quadrant. (Credit: AI IQ)

                                      The AI cost-performance chart that enterprise buyers actually need to see

                                      Perhaps the most practically useful chart on the site is not the bell curve but the IQ vs. Effective Cost scatter plot. It maps each model’s estimated IQ against an «effective cost» metric — defined as the token cost for a task using 2 million input tokens and 1 million output tokens, multiplied by a usage efficiency factor.

                                      The chart reveals a familiar pattern in enterprise technology: the best models are not always the best value. GPT-5.5 and Opus 4.7 sit in the upper-left corner — high IQ, high cost, with effective per-task costs north of $30 and $50 respectively. Meanwhile, models like GPT-5.4-mini, DeepSeek-V3.2, and MiniMax-M2.7 occupy a sweet spot in the middle: respectable IQ scores between 112 and 120, at effective costs ranging from roughly $1 to $5 per task. At the cheapest extreme, GPT-oss-20b (an open-source OpenAI model) appears near $0.20 effective cost with an IQ around 107 — potentially the most economical option for bulk classification or extraction workloads.

                                      The site also offers a 3D visualization mapping IQ, EQ, and effective cost simultaneously. A dashed line running through the cube points toward the ideal: higher IQ, higher EQ, and lower cost. Models near the «green end» of that axis are stronger all-around deals; those near the «red end» sacrifice capability, cost efficiency, or both. For CIOs staring at API invoices, the implication is clear: the intelligence gap between a $50 model and a $3 model has narrowed enough that routing — using expensive models for hard problems and cheap ones for everything else — is no longer optional. It is the dominant architecture for serious AI deployments.

                                      Critics say AI’s «jagged» capabilities make a single IQ score dangerously misleading

                                      The loudest objection to AI IQ is philosophical, and it cuts deep. Critics argue that collapsing a model’s uneven capabilities into a single score obscures more than it reveals.

                                      «IQ as a proxy is fading — we’re seeing reasoning density spikes that don’t map to g-factor,» posted Zaya, a technology commentator, on X. «GPT-5.5 already hit saturation on MMLU-Pro, but still fails ClockBench 50% of the time.»

                                      That observation touches on what AI researchers call the «jaggedness» problem: large language models often exhibit wildly uneven capabilities, excelling at graduate-level physics while failing at tasks a child could do. A composite score can paper over those gaps.

                                      Pressureangle, another X user, posted a more granular critique, calling out «complete lack of transparency» and arguing the site never fully discloses how its calibration curves were created or validated. In fairness, AI IQ does list its 12 benchmarks and shows the shape of each calibration curve in its methodology modal. But the raw data and precise mathematical transformations are not published as open datasets — a gap that matters to researchers accustomed to fully reproducible methods.

                                      Others questioned the premise itself. «As useless as human IQ testing,» wrote haashim on X. Shubham Sharma, an AI and technology writer, offered a constructive alternative: «Why not having the Models take an official (MENSA-Grade) test? Wouldn’t this be the most accurate and most ‘human-comparable’ way to benchmark intelligence?» That approach already exists through TrackingAI, which administers the Mensa Norway IQ test to language models. But Mensa-style tests measure only abstract pattern recognition, while AI IQ attempts a broader composite across coding, mathematics, and academic reasoning. As Visual Capitalist noted, «an IQ-style benchmark captures only one slice of capability.» Each approach has tradeoffs — and neither has won the argument yet.

                                      The real race isn’t for the highest score — it’s for the smartest model stack

                                      For all the debate about methodology, the most important signal in AI IQ’s data may not be any single model’s score. It is the shape of the market the charts reveal.

                                      There are now more than 50 frontier-class models available through APIs, from at least 14 major providers spanning the United States, China, and Europe. Each provider publishes its own benchmarks, often cherry-picked to showcase strengths. The result is a Tower of Babel where no two companies measure the same thing in the same way. Academic research has highlighted that «most benchmarks introduce bias by focusing on a particular type of domain,» and the Frontier IQ Over Time chart on AI IQ shows just how fast the targets are moving: in October 2023, GPT-4-turbo sat near an estimated IQ of 75. By early 2026, the top models were brushing 135 — roughly 60 points of improvement in 30 months.

                                      That pace raises a fundamental question about whether any scoring system can keep up. The site compresses ceilings for saturated benchmarks, but as models continue to max out even the hardest tests — ARC-AGI-2, FrontierMath Tier 4, Humanity’s Last Exam — the framework will face the same ceiling effects that have plagued every AI evaluation before it. Connor Forsyth pointed to this dynamic on X: «ARC AGI 3 disagrees,» he wrote, referencing a next-generation benchmark that may already be undermining current scores.

                                      AI IQ is not perfect. Its methodology is partially opaque. Its IQ metaphor can mislead. And its creator acknowledges known biases while likely missing others. But the alternative — wading through dozens of provider-specific benchmark tables, each using different test suites and scoring conventions — is worse. The site offers enterprise buyers something genuinely scarce: a single framework for comparing models across providers, dimensions, and price points, updated regularly, with enough nuance to show that the right answer to «which model is best?» is almost always «it depends on the task.»

                                      As Debdoot Ghosh mused on X after viewing the charts: «Now a human’s role is just to orchestrate?«

                                      Maybe. But if the AI IQ data shows anything clearly, it is that orchestration — knowing which model to deploy, when, and at what price — has become its own form of intelligence. And for that, there is no benchmark yet.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      For decades, the IQ test has been one of the most familiar — and most contested — yardsticks for human intelligence. Now, a startup project called AI IQ is applying the same metaphor to artificial intelligence, assigning estimated intelligence quotients to more than 50 of the world’s most powerful language models and plotting them on a standard bell curve.

                                      The result is a set of interactive visualizations at aiiq.org that have ricocheted across social media in the past week, drawing praise from enterprise technologists who say the charts make an impossibly complex market legible — and sharp criticism from researchers and commentators who warn the entire framework is misleading.

                                      «This is super useful,» wrote Thibaut Mélen, a technology commentator, on X. «Much easier to understand model progress when it’s mapped like this instead of another giant leaderboard table.»

                                      Brian Vellmure, a business strategist, offered a similar endorsement: «This is helpful. Anecdotally tracks with personal experience.»

                                      But the backlash arrived just as quickly. «It’s nonsense. AI is far too jagged. The map is not the territory,» posted AI Deeply, an artificial intelligence commentary account, crystallizing a worry shared by many researchers: that reducing a language model’s sprawling, uneven capabilities to a single number creates a dangerous illusion of precision.

                                      More than 50 AI language models, plotted on a standard IQ bell curve by the site AI IQ. The most capable models crowd the right tail of the distribution. (Credit: AI IQ)

                                      Twelve benchmarks, four dimensions, and one controversial number: how AI IQ actually works

                                      AI IQ was created by Ryan Shea, an engineer, entrepreneur, and angel investor best known as a co-founder of the blockchain platform Stacks. Shea also co-founded Voterbase and has invested in the early stages of several unicorns, including OpenSea, Lattice, Anchorage, and Mercury. He holds a Bachelor of Science in Mechanical Engineering from Princeton University.

                                      The site’s methodology rests on a deceptively simple formula. AI IQ groups 12 benchmarks into four reasoning dimensions: abstract, mathematical, programmatic, and academic. The composite IQ is a straight average of those four dimension scores: IQ = ¼ (IQ_Abstract + IQ_Math + IQ_Prog + IQ_Acad).

                                      The abstract reasoning dimension draws from ARC-AGI-1 and ARC-AGI-2, the notoriously difficult pattern-recognition benchmarks designed to test general fluid intelligence. Mathematical reasoning includes FrontierMath (Tiers 1–3 and Tier 4), AIME, and ProofBench. Programmatic reasoning uses Terminal-Bench 2.0, SWE-Bench Verified, and SciCode. Academic reasoning pulls from Humanity’s Last Exam, CritPt, and GPQA Diamond.

                                      Each raw benchmark score gets mapped to an implied IQ through what the site describes as «hand-calibrated difficulty curves.» Crucially, the methodology compresses ceilings for benchmarks considered easier or more susceptible to data contamination, preventing them from inflating scores above 100. Harder, less gameable benchmarks retain higher ceilings. The system also handles missing data conservatively: models need scores on at least two of the four dimensions to receive a derived IQ, and when benchmarks are absent, the pipeline deliberately pulls scores down rather than up. The site states that «every derived IQ averages all four dimensions, so missing coverage cannot make a model look better by omission.»

                                      OpenAI leads the bell curve, but the gap between the top AI models has never been smaller

                                      As of mid-May 2026, the AI IQ charts tell a story of rapid convergence at the top of the frontier — and widening diversity in the tiers below.

                                      According to the Frontier IQ Over Time chart, GPT-5.5 from OpenAI currently sits at the peak of the bell curve, with an estimated IQ near 136 — the highest of any model tracked. It is closely followed by GPT-5.4 (approximately 131), Opus 4.7 from Anthropic (approximately 132), and Opus 4.6 (approximately 129). Google’s Gemini 3.1 Pro lands near 131, making the top cluster extraordinarily tight.

                                      That compression is not unique to AI IQ’s framework. Visual Capitalist, drawing from a separate Mensa-based ranking by TrackingAI, recently observed the same dynamic, noting that «the biggest takeaway is how compressed the top of the leaderboard has become.» On that scale, Grok-4.20 Expert Mode and GPT 5.4 Pro tied at 145, with Gemini 3.1 Pro at 141.

                                      Below the frontier cluster, the AI IQ charts show a crowded midfield. Models from Chinese labs — Kimi K2.6, GLM-5, DeepSeek-V3.2, Qwen3.6, MiniMax-M2.7 — bunch between roughly 112 and 118, making the cost-performance tier increasingly competitive for enterprise buyers who don’t need the absolute best model for every task. One X user, ovsky, noted that the data «confirms experience with sonnet 4.6 being an absolute workhorse as opposed to opus 4.5» — pointing to the way the charts can validate practitioner intuitions that headline rankings often miss.

                                      aiiq-frontier-iq-over-time-2026-05-13

                                      The trajectory of frontier AI models from October 2023 to mid-2026, as tracked by AI IQ. Provider-colored step-lines connect each lab’s flagship releases, showing roughly 60 points of estimated IQ improvement in 30 months. (Credit: AI IQ)

                                      Why emotional intelligence scores are becoming the new battleground in AI model rankings

                                      What distinguishes AI IQ from most other benchmarking efforts is its inclusion of an «EQ» — emotional intelligence — score. The site maps each model’s EQ-Bench 3 Elo score and Arena Elo score to an estimated EQ using calibrated piecewise-linear scales, then takes a 50/50 weighted composite of the two.

                                      The EQ scores produce a meaningfully different ranking than IQ alone. On the IQ vs. EQ scatter plot, Anthropic’s Opus 4.7 leads on EQ with a score near 132, pushing it into the upper-right quadrant — the most desirable position, signaling both high cognitive and high emotional intelligence. OpenAI’s GPT-5.5 and GPT-5.4 cluster in the high-IQ zone but lag slightly on EQ. Google’s Gemini 3.1 Pro sits in a strong middle position on both axes.

                                      One notable methodological choice has drawn attention: EQ-Bench 3 is judged by Claude, an Anthropic model, which the site acknowledges «creates potential scoring bias in favor of Anthropic models.» To correct for this, AI IQ subtracts a 200-point Elo penalty from the EQ-Bench component for all Anthropic models before mapping to implied EQ. The Arena component is unaffected since it uses human judges. That self-correction is unusual in the benchmarking world, and it suggests Shea is aware of the methodological minefield he has entered. Still, the EQ dimension captures something IQ alone cannot: the growing importance of conversational quality, collaboration, and trust in models deployed for user-facing work.

                                      aiiq-iq-vs-eq-2026-05-13

                                      Plotting IQ against EQ reveals that the smartest models aren’t always the most emotionally intelligent. Anthropic’s Opus 4.7 dominates the upper-right quadrant. (Credit: AI IQ)

                                      The AI cost-performance chart that enterprise buyers actually need to see

                                      Perhaps the most practically useful chart on the site is not the bell curve but the IQ vs. Effective Cost scatter plot. It maps each model’s estimated IQ against an «effective cost» metric — defined as the token cost for a task using 2 million input tokens and 1 million output tokens, multiplied by a usage efficiency factor.

                                      The chart reveals a familiar pattern in enterprise technology: the best models are not always the best value. GPT-5.5 and Opus 4.7 sit in the upper-left corner — high IQ, high cost, with effective per-task costs north of $30 and $50 respectively. Meanwhile, models like GPT-5.4-mini, DeepSeek-V3.2, and MiniMax-M2.7 occupy a sweet spot in the middle: respectable IQ scores between 112 and 120, at effective costs ranging from roughly $1 to $5 per task. At the cheapest extreme, GPT-oss-20b (an open-source OpenAI model) appears near $0.20 effective cost with an IQ around 107 — potentially the most economical option for bulk classification or extraction workloads.

                                      The site also offers a 3D visualization mapping IQ, EQ, and effective cost simultaneously. A dashed line running through the cube points toward the ideal: higher IQ, higher EQ, and lower cost. Models near the «green end» of that axis are stronger all-around deals; those near the «red end» sacrifice capability, cost efficiency, or both. For CIOs staring at API invoices, the implication is clear: the intelligence gap between a $50 model and a $3 model has narrowed enough that routing — using expensive models for hard problems and cheap ones for everything else — is no longer optional. It is the dominant architecture for serious AI deployments.

                                      Critics say AI’s «jagged» capabilities make a single IQ score dangerously misleading

                                      The loudest objection to AI IQ is philosophical, and it cuts deep. Critics argue that collapsing a model’s uneven capabilities into a single score obscures more than it reveals.

                                      «IQ as a proxy is fading — we’re seeing reasoning density spikes that don’t map to g-factor,» posted Zaya, a technology commentator, on X. «GPT-5.5 already hit saturation on MMLU-Pro, but still fails ClockBench 50% of the time.»

                                      That observation touches on what AI researchers call the «jaggedness» problem: large language models often exhibit wildly uneven capabilities, excelling at graduate-level physics while failing at tasks a child could do. A composite score can paper over those gaps.

                                      Pressureangle, another X user, posted a more granular critique, calling out «complete lack of transparency» and arguing the site never fully discloses how its calibration curves were created or validated. In fairness, AI IQ does list its 12 benchmarks and shows the shape of each calibration curve in its methodology modal. But the raw data and precise mathematical transformations are not published as open datasets — a gap that matters to researchers accustomed to fully reproducible methods.

                                      Others questioned the premise itself. «As useless as human IQ testing,» wrote haashim on X. Shubham Sharma, an AI and technology writer, offered a constructive alternative: «Why not having the Models take an official (MENSA-Grade) test? Wouldn’t this be the most accurate and most ‘human-comparable’ way to benchmark intelligence?» That approach already exists through TrackingAI, which administers the Mensa Norway IQ test to language models. But Mensa-style tests measure only abstract pattern recognition, while AI IQ attempts a broader composite across coding, mathematics, and academic reasoning. As Visual Capitalist noted, «an IQ-style benchmark captures only one slice of capability.» Each approach has tradeoffs — and neither has won the argument yet.

                                      The real race isn’t for the highest score — it’s for the smartest model stack

                                      For all the debate about methodology, the most important signal in AI IQ’s data may not be any single model’s score. It is the shape of the market the charts reveal.

                                      There are now more than 50 frontier-class models available through APIs, from at least 14 major providers spanning the United States, China, and Europe. Each provider publishes its own benchmarks, often cherry-picked to showcase strengths. The result is a Tower of Babel where no two companies measure the same thing in the same way. Academic research has highlighted that «most benchmarks introduce bias by focusing on a particular type of domain,» and the Frontier IQ Over Time chart on AI IQ shows just how fast the targets are moving: in October 2023, GPT-4-turbo sat near an estimated IQ of 75. By early 2026, the top models were brushing 135 — roughly 60 points of improvement in 30 months.

                                      That pace raises a fundamental question about whether any scoring system can keep up. The site compresses ceilings for saturated benchmarks, but as models continue to max out even the hardest tests — ARC-AGI-2, FrontierMath Tier 4, Humanity’s Last Exam — the framework will face the same ceiling effects that have plagued every AI evaluation before it. Connor Forsyth pointed to this dynamic on X: «ARC AGI 3 disagrees,» he wrote, referencing a next-generation benchmark that may already be undermining current scores.

                                      AI IQ is not perfect. Its methodology is partially opaque. Its IQ metaphor can mislead. And its creator acknowledges known biases while likely missing others. But the alternative — wading through dozens of provider-specific benchmark tables, each using different test suites and scoring conventions — is worse. The site offers enterprise buyers something genuinely scarce: a single framework for comparing models across providers, dimensions, and price points, updated regularly, with enough nuance to show that the right answer to «which model is best?» is almost always «it depends on the task.»

                                      As Debdoot Ghosh mused on X after viewing the charts: «Now a human’s role is just to orchestrate?«

                                      Maybe. But if the AI IQ data shows anything clearly, it is that orchestration — knowing which model to deploy, when, and at what price — has become its own form of intelligence. And for that, there is no benchmark yet.

                                      ¡No te pierdas las noticias destacadas!

                                      Suscríbete y recibe las historias más importantes del día.

                                      Al suscribirte aceptas nuestros términos y condiciones y política de privacidad.

                                      For decades, the IQ test has been one of the most familiar — and most contested — yardsticks for human intelligence. Now, a startup project called AI IQ is applying the same metaphor to artificial intelligence, assigning estimated intelligence quotients to more than 50 of the world’s most powerful language models and plotting them on a standard bell curve.

                                      The result is a set of interactive visualizations at aiiq.org that have ricocheted across social media in the past week, drawing praise from enterprise technologists who say the charts make an impossibly complex market legible — and sharp criticism from researchers and commentators who warn the entire framework is misleading.

                                      «This is super useful,» wrote Thibaut Mélen, a technology commentator, on X. «Much easier to understand model progress when it’s mapped like this instead of another giant leaderboard table.»

                                      Brian Vellmure, a business strategist, offered a similar endorsement: «This is helpful. Anecdotally tracks with personal experience.»

                                      But the backlash arrived just as quickly. «It’s nonsense. AI is far too jagged. The map is not the territory,» posted AI Deeply, an artificial intelligence commentary account, crystallizing a worry shared by many researchers: that reducing a language model’s sprawling, uneven capabilities to a single number creates a dangerous illusion of precision.

                                      More than 50 AI language models, plotted on a standard IQ bell curve by the site AI IQ. The most capable models crowd the right tail of the distribution. (Credit: AI IQ)

                                      Twelve benchmarks, four dimensions, and one controversial number: how AI IQ actually works

                                      AI IQ was created by Ryan Shea, an engineer, entrepreneur, and angel investor best known as a co-founder of the blockchain platform Stacks. Shea also co-founded Voterbase and has invested in the early stages of several unicorns, including OpenSea, Lattice, Anchorage, and Mercury. He holds a Bachelor of Science in Mechanical Engineering from Princeton University.

                                      The site’s methodology rests on a deceptively simple formula. AI IQ groups 12 benchmarks into four reasoning dimensions: abstract, mathematical, programmatic, and academic. The composite IQ is a straight average of those four dimension scores: IQ = ¼ (IQ_Abstract + IQ_Math + IQ_Prog + IQ_Acad).

                                      The abstract reasoning dimension draws from ARC-AGI-1 and ARC-AGI-2, the notoriously difficult pattern-recognition benchmarks designed to test general fluid intelligence. Mathematical reasoning includes FrontierMath (Tiers 1–3 and Tier 4), AIME, and ProofBench. Programmatic reasoning uses Terminal-Bench 2.0, SWE-Bench Verified, and SciCode. Academic reasoning pulls from Humanity’s Last Exam, CritPt, and GPQA Diamond.

                                      Each raw benchmark score gets mapped to an implied IQ through what the site describes as «hand-calibrated difficulty curves.» Crucially, the methodology compresses ceilings for benchmarks considered easier or more susceptible to data contamination, preventing them from inflating scores above 100. Harder, less gameable benchmarks retain higher ceilings. The system also handles missing data conservatively: models need scores on at least two of the four dimensions to receive a derived IQ, and when benchmarks are absent, the pipeline deliberately pulls scores down rather than up. The site states that «every derived IQ averages all four dimensions, so missing coverage cannot make a model look better by omission.»

                                      OpenAI leads the bell curve, but the gap between the top AI models has never been smaller

                                      As of mid-May 2026, the AI IQ charts tell a story of rapid convergence at the top of the frontier — and widening diversity in the tiers below.

                                      According to the Frontier IQ Over Time chart, GPT-5.5 from OpenAI currently sits at the peak of the bell curve, with an estimated IQ near 136 — the highest of any model tracked. It is closely followed by GPT-5.4 (approximately 131), Opus 4.7 from Anthropic (approximately 132), and Opus 4.6 (approximately 129). Google’s Gemini 3.1 Pro lands near 131, making the top cluster extraordinarily tight.

                                      That compression is not unique to AI IQ’s framework. Visual Capitalist, drawing from a separate Mensa-based ranking by TrackingAI, recently observed the same dynamic, noting that «the biggest takeaway is how compressed the top of the leaderboard has become.» On that scale, Grok-4.20 Expert Mode and GPT 5.4 Pro tied at 145, with Gemini 3.1 Pro at 141.

                                      Below the frontier cluster, the AI IQ charts show a crowded midfield. Models from Chinese labs — Kimi K2.6, GLM-5, DeepSeek-V3.2, Qwen3.6, MiniMax-M2.7 — bunch between roughly 112 and 118, making the cost-performance tier increasingly competitive for enterprise buyers who don’t need the absolute best model for every task. One X user, ovsky, noted that the data «confirms experience with sonnet 4.6 being an absolute workhorse as opposed to opus 4.5» — pointing to the way the charts can validate practitioner intuitions that headline rankings often miss.

                                      aiiq-frontier-iq-over-time-2026-05-13

                                      The trajectory of frontier AI models from October 2023 to mid-2026, as tracked by AI IQ. Provider-colored step-lines connect each lab’s flagship releases, showing roughly 60 points of estimated IQ improvement in 30 months. (Credit: AI IQ)

                                      Why emotional intelligence scores are becoming the new battleground in AI model rankings

                                      What distinguishes AI IQ from most other benchmarking efforts is its inclusion of an «EQ» — emotional intelligence — score. The site maps each model’s EQ-Bench 3 Elo score and Arena Elo score to an estimated EQ using calibrated piecewise-linear scales, then takes a 50/50 weighted composite of the two.

                                      The EQ scores produce a meaningfully different ranking than IQ alone. On the IQ vs. EQ scatter plot, Anthropic’s Opus 4.7 leads on EQ with a score near 132, pushing it into the upper-right quadrant — the most desirable position, signaling both high cognitive and high emotional intelligence. OpenAI’s GPT-5.5 and GPT-5.4 cluster in the high-IQ zone but lag slightly on EQ. Google’s Gemini 3.1 Pro sits in a strong middle position on both axes.

                                      One notable methodological choice has drawn attention: EQ-Bench 3 is judged by Claude, an Anthropic model, which the site acknowledges «creates potential scoring bias in favor of Anthropic models.» To correct for this, AI IQ subtracts a 200-point Elo penalty from the EQ-Bench component for all Anthropic models before mapping to implied EQ. The Arena component is unaffected since it uses human judges. That self-correction is unusual in the benchmarking world, and it suggests Shea is aware of the methodological minefield he has entered. Still, the EQ dimension captures something IQ alone cannot: the growing importance of conversational quality, collaboration, and trust in models deployed for user-facing work.

                                      aiiq-iq-vs-eq-2026-05-13

                                      Plotting IQ against EQ reveals that the smartest models aren’t always the most emotionally intelligent. Anthropic’s Opus 4.7 dominates the upper-right quadrant. (Credit: AI IQ)

                                      The AI cost-performance chart that enterprise buyers actually need to see

                                      Perhaps the most practically useful chart on the site is not the bell curve but the IQ vs. Effective Cost scatter plot. It maps each model’s estimated IQ against an «effective cost» metric — defined as the token cost for a task using 2 million input tokens and 1 million output tokens, multiplied by a usage efficiency factor.

                                      The chart reveals a familiar pattern in enterprise technology: the best models are not always the best value. GPT-5.5 and Opus 4.7 sit in the upper-left corner — high IQ, high cost, with effective per-task costs north of $30 and $50 respectively. Meanwhile, models like GPT-5.4-mini, DeepSeek-V3.2, and MiniMax-M2.7 occupy a sweet spot in the middle: respectable IQ scores between 112 and 120, at effective costs ranging from roughly $1 to $5 per task. At the cheapest extreme, GPT-oss-20b (an open-source OpenAI model) appears near $0.20 effective cost with an IQ around 107 — potentially the most economical option for bulk classification or extraction workloads.

                                      The site also offers a 3D visualization mapping IQ, EQ, and effective cost simultaneously. A dashed line running through the cube points toward the ideal: higher IQ, higher EQ, and lower cost. Models near the «green end» of that axis are stronger all-around deals; those near the «red end» sacrifice capability, cost efficiency, or both. For CIOs staring at API invoices, the implication is clear: the intelligence gap between a $50 model and a $3 model has narrowed enough that routing — using expensive models for hard problems and cheap ones for everything else — is no longer optional. It is the dominant architecture for serious AI deployments.

                                      Critics say AI’s «jagged» capabilities make a single IQ score dangerously misleading

                                      The loudest objection to AI IQ is philosophical, and it cuts deep. Critics argue that collapsing a model’s uneven capabilities into a single score obscures more than it reveals.

                                      «IQ as a proxy is fading — we’re seeing reasoning density spikes that don’t map to g-factor,» posted Zaya, a technology commentator, on X. «GPT-5.5 already hit saturation on MMLU-Pro, but still fails ClockBench 50% of the time.»

                                      That observation touches on what AI researchers call the «jaggedness» problem: large language models often exhibit wildly uneven capabilities, excelling at graduate-level physics while failing at tasks a child could do. A composite score can paper over those gaps.

                                      Pressureangle, another X user, posted a more granular critique, calling out «complete lack of transparency» and arguing the site never fully discloses how its calibration curves were created or validated. In fairness, AI IQ does list its 12 benchmarks and shows the shape of each calibration curve in its methodology modal. But the raw data and precise mathematical transformations are not published as open datasets — a gap that matters to researchers accustomed to fully reproducible methods.

                                      Others questioned the premise itself. «As useless as human IQ testing,» wrote haashim on X. Shubham Sharma, an AI and technology writer, offered a constructive alternative: «Why not having the Models take an official (MENSA-Grade) test? Wouldn’t this be the most accurate and most ‘human-comparable’ way to benchmark intelligence?» That approach already exists through TrackingAI, which administers the Mensa Norway IQ test to language models. But Mensa-style tests measure only abstract pattern recognition, while AI IQ attempts a broader composite across coding, mathematics, and academic reasoning. As Visual Capitalist noted, «an IQ-style benchmark captures only one slice of capability.» Each approach has tradeoffs — and neither has won the argument yet.

                                      The real race isn’t for the highest score — it’s for the smartest model stack

                                      For all the debate about methodology, the most important signal in AI IQ’s data may not be any single model’s score. It is the shape of the market the charts reveal.

                                      There are now more than 50 frontier-class models available through APIs, from at least 14 major providers spanning the United States, China, and Europe. Each provider publishes its own benchmarks, often cherry-picked to showcase strengths. The result is a Tower of Babel where no two companies measure the same thing in the same way. Academic research has highlighted that «most benchmarks introduce bias by focusing on a particular type of domain,» and the Frontier IQ Over Time chart on AI IQ shows just how fast the targets are moving: in October 2023, GPT-4-turbo sat near an estimated IQ of 75. By early 2026, the top models were brushing 135 — roughly 60 points of improvement in 30 months.

                                      That pace raises a fundamental question about whether any scoring system can keep up. The site compresses ceilings for saturated benchmarks, but as models continue to max out even the hardest tests — ARC-AGI-2, FrontierMath Tier 4, Humanity’s Last Exam — the framework will face the same ceiling effects that have plagued every AI evaluation before it. Connor Forsyth pointed to this dynamic on X: «ARC AGI 3 disagrees,» he wrote, referencing a next-generation benchmark that may already be undermining current scores.

                                      AI IQ is not perfect. Its methodology is partially opaque. Its IQ metaphor can mislead. And its creator acknowledges known biases while likely missing others. But the alternative — wading through dozens of provider-specific benchmark tables, each using different test suites and scoring conventions — is worse. The site offers enterprise buyers something genuinely scarce: a single framework for comparing models across providers, dimensions, and price points, updated regularly, with enough nuance to show that the right answer to «which model is best?» is almost always «it depends on the task.»

                                      As Debdoot Ghosh mused on X after viewing the charts: «Now a human’s role is just to orchestrate?«

                                      Maybe. But if the AI IQ data shows anything clearly, it is that orchestration — knowing which model to deploy, when, and at what price — has become its own form of intelligence. And for that, there is no benchmark yet.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      For decades, the IQ test has been one of the most familiar — and most contested — yardsticks for human intelligence. Now, a startup project called AI IQ is applying the same metaphor to artificial intelligence, assigning estimated intelligence quotients to more than 50 of the world’s most powerful language models and plotting them on a standard bell curve.

                                      The result is a set of interactive visualizations at aiiq.org that have ricocheted across social media in the past week, drawing praise from enterprise technologists who say the charts make an impossibly complex market legible — and sharp criticism from researchers and commentators who warn the entire framework is misleading.

                                      «This is super useful,» wrote Thibaut Mélen, a technology commentator, on X. «Much easier to understand model progress when it’s mapped like this instead of another giant leaderboard table.»

                                      Brian Vellmure, a business strategist, offered a similar endorsement: «This is helpful. Anecdotally tracks with personal experience.»

                                      But the backlash arrived just as quickly. «It’s nonsense. AI is far too jagged. The map is not the territory,» posted AI Deeply, an artificial intelligence commentary account, crystallizing a worry shared by many researchers: that reducing a language model’s sprawling, uneven capabilities to a single number creates a dangerous illusion of precision.

                                      More than 50 AI language models, plotted on a standard IQ bell curve by the site AI IQ. The most capable models crowd the right tail of the distribution. (Credit: AI IQ)

                                      Twelve benchmarks, four dimensions, and one controversial number: how AI IQ actually works

                                      AI IQ was created by Ryan Shea, an engineer, entrepreneur, and angel investor best known as a co-founder of the blockchain platform Stacks. Shea also co-founded Voterbase and has invested in the early stages of several unicorns, including OpenSea, Lattice, Anchorage, and Mercury. He holds a Bachelor of Science in Mechanical Engineering from Princeton University.

                                      The site’s methodology rests on a deceptively simple formula. AI IQ groups 12 benchmarks into four reasoning dimensions: abstract, mathematical, programmatic, and academic. The composite IQ is a straight average of those four dimension scores: IQ = ¼ (IQ_Abstract + IQ_Math + IQ_Prog + IQ_Acad).

                                      The abstract reasoning dimension draws from ARC-AGI-1 and ARC-AGI-2, the notoriously difficult pattern-recognition benchmarks designed to test general fluid intelligence. Mathematical reasoning includes FrontierMath (Tiers 1–3 and Tier 4), AIME, and ProofBench. Programmatic reasoning uses Terminal-Bench 2.0, SWE-Bench Verified, and SciCode. Academic reasoning pulls from Humanity’s Last Exam, CritPt, and GPQA Diamond.

                                      Each raw benchmark score gets mapped to an implied IQ through what the site describes as «hand-calibrated difficulty curves.» Crucially, the methodology compresses ceilings for benchmarks considered easier or more susceptible to data contamination, preventing them from inflating scores above 100. Harder, less gameable benchmarks retain higher ceilings. The system also handles missing data conservatively: models need scores on at least two of the four dimensions to receive a derived IQ, and when benchmarks are absent, the pipeline deliberately pulls scores down rather than up. The site states that «every derived IQ averages all four dimensions, so missing coverage cannot make a model look better by omission.»

                                      OpenAI leads the bell curve, but the gap between the top AI models has never been smaller

                                      As of mid-May 2026, the AI IQ charts tell a story of rapid convergence at the top of the frontier — and widening diversity in the tiers below.

                                      According to the Frontier IQ Over Time chart, GPT-5.5 from OpenAI currently sits at the peak of the bell curve, with an estimated IQ near 136 — the highest of any model tracked. It is closely followed by GPT-5.4 (approximately 131), Opus 4.7 from Anthropic (approximately 132), and Opus 4.6 (approximately 129). Google’s Gemini 3.1 Pro lands near 131, making the top cluster extraordinarily tight.

                                      That compression is not unique to AI IQ’s framework. Visual Capitalist, drawing from a separate Mensa-based ranking by TrackingAI, recently observed the same dynamic, noting that «the biggest takeaway is how compressed the top of the leaderboard has become.» On that scale, Grok-4.20 Expert Mode and GPT 5.4 Pro tied at 145, with Gemini 3.1 Pro at 141.

                                      Below the frontier cluster, the AI IQ charts show a crowded midfield. Models from Chinese labs — Kimi K2.6, GLM-5, DeepSeek-V3.2, Qwen3.6, MiniMax-M2.7 — bunch between roughly 112 and 118, making the cost-performance tier increasingly competitive for enterprise buyers who don’t need the absolute best model for every task. One X user, ovsky, noted that the data «confirms experience with sonnet 4.6 being an absolute workhorse as opposed to opus 4.5» — pointing to the way the charts can validate practitioner intuitions that headline rankings often miss.

                                      aiiq-frontier-iq-over-time-2026-05-13

                                      The trajectory of frontier AI models from October 2023 to mid-2026, as tracked by AI IQ. Provider-colored step-lines connect each lab’s flagship releases, showing roughly 60 points of estimated IQ improvement in 30 months. (Credit: AI IQ)

                                      Why emotional intelligence scores are becoming the new battleground in AI model rankings

                                      What distinguishes AI IQ from most other benchmarking efforts is its inclusion of an «EQ» — emotional intelligence — score. The site maps each model’s EQ-Bench 3 Elo score and Arena Elo score to an estimated EQ using calibrated piecewise-linear scales, then takes a 50/50 weighted composite of the two.

                                      The EQ scores produce a meaningfully different ranking than IQ alone. On the IQ vs. EQ scatter plot, Anthropic’s Opus 4.7 leads on EQ with a score near 132, pushing it into the upper-right quadrant — the most desirable position, signaling both high cognitive and high emotional intelligence. OpenAI’s GPT-5.5 and GPT-5.4 cluster in the high-IQ zone but lag slightly on EQ. Google’s Gemini 3.1 Pro sits in a strong middle position on both axes.

                                      One notable methodological choice has drawn attention: EQ-Bench 3 is judged by Claude, an Anthropic model, which the site acknowledges «creates potential scoring bias in favor of Anthropic models.» To correct for this, AI IQ subtracts a 200-point Elo penalty from the EQ-Bench component for all Anthropic models before mapping to implied EQ. The Arena component is unaffected since it uses human judges. That self-correction is unusual in the benchmarking world, and it suggests Shea is aware of the methodological minefield he has entered. Still, the EQ dimension captures something IQ alone cannot: the growing importance of conversational quality, collaboration, and trust in models deployed for user-facing work.

                                      aiiq-iq-vs-eq-2026-05-13

                                      Plotting IQ against EQ reveals that the smartest models aren’t always the most emotionally intelligent. Anthropic’s Opus 4.7 dominates the upper-right quadrant. (Credit: AI IQ)

                                      The AI cost-performance chart that enterprise buyers actually need to see

                                      Perhaps the most practically useful chart on the site is not the bell curve but the IQ vs. Effective Cost scatter plot. It maps each model’s estimated IQ against an «effective cost» metric — defined as the token cost for a task using 2 million input tokens and 1 million output tokens, multiplied by a usage efficiency factor.

                                      The chart reveals a familiar pattern in enterprise technology: the best models are not always the best value. GPT-5.5 and Opus 4.7 sit in the upper-left corner — high IQ, high cost, with effective per-task costs north of $30 and $50 respectively. Meanwhile, models like GPT-5.4-mini, DeepSeek-V3.2, and MiniMax-M2.7 occupy a sweet spot in the middle: respectable IQ scores between 112 and 120, at effective costs ranging from roughly $1 to $5 per task. At the cheapest extreme, GPT-oss-20b (an open-source OpenAI model) appears near $0.20 effective cost with an IQ around 107 — potentially the most economical option for bulk classification or extraction workloads.

                                      The site also offers a 3D visualization mapping IQ, EQ, and effective cost simultaneously. A dashed line running through the cube points toward the ideal: higher IQ, higher EQ, and lower cost. Models near the «green end» of that axis are stronger all-around deals; those near the «red end» sacrifice capability, cost efficiency, or both. For CIOs staring at API invoices, the implication is clear: the intelligence gap between a $50 model and a $3 model has narrowed enough that routing — using expensive models for hard problems and cheap ones for everything else — is no longer optional. It is the dominant architecture for serious AI deployments.

                                      Critics say AI’s «jagged» capabilities make a single IQ score dangerously misleading

                                      The loudest objection to AI IQ is philosophical, and it cuts deep. Critics argue that collapsing a model’s uneven capabilities into a single score obscures more than it reveals.

                                      «IQ as a proxy is fading — we’re seeing reasoning density spikes that don’t map to g-factor,» posted Zaya, a technology commentator, on X. «GPT-5.5 already hit saturation on MMLU-Pro, but still fails ClockBench 50% of the time.»

                                      That observation touches on what AI researchers call the «jaggedness» problem: large language models often exhibit wildly uneven capabilities, excelling at graduate-level physics while failing at tasks a child could do. A composite score can paper over those gaps.

                                      Pressureangle, another X user, posted a more granular critique, calling out «complete lack of transparency» and arguing the site never fully discloses how its calibration curves were created or validated. In fairness, AI IQ does list its 12 benchmarks and shows the shape of each calibration curve in its methodology modal. But the raw data and precise mathematical transformations are not published as open datasets — a gap that matters to researchers accustomed to fully reproducible methods.

                                      Others questioned the premise itself. «As useless as human IQ testing,» wrote haashim on X. Shubham Sharma, an AI and technology writer, offered a constructive alternative: «Why not having the Models take an official (MENSA-Grade) test? Wouldn’t this be the most accurate and most ‘human-comparable’ way to benchmark intelligence?» That approach already exists through TrackingAI, which administers the Mensa Norway IQ test to language models. But Mensa-style tests measure only abstract pattern recognition, while AI IQ attempts a broader composite across coding, mathematics, and academic reasoning. As Visual Capitalist noted, «an IQ-style benchmark captures only one slice of capability.» Each approach has tradeoffs — and neither has won the argument yet.

                                      The real race isn’t for the highest score — it’s for the smartest model stack

                                      For all the debate about methodology, the most important signal in AI IQ’s data may not be any single model’s score. It is the shape of the market the charts reveal.

                                      There are now more than 50 frontier-class models available through APIs, from at least 14 major providers spanning the United States, China, and Europe. Each provider publishes its own benchmarks, often cherry-picked to showcase strengths. The result is a Tower of Babel where no two companies measure the same thing in the same way. Academic research has highlighted that «most benchmarks introduce bias by focusing on a particular type of domain,» and the Frontier IQ Over Time chart on AI IQ shows just how fast the targets are moving: in October 2023, GPT-4-turbo sat near an estimated IQ of 75. By early 2026, the top models were brushing 135 — roughly 60 points of improvement in 30 months.

                                      That pace raises a fundamental question about whether any scoring system can keep up. The site compresses ceilings for saturated benchmarks, but as models continue to max out even the hardest tests — ARC-AGI-2, FrontierMath Tier 4, Humanity’s Last Exam — the framework will face the same ceiling effects that have plagued every AI evaluation before it. Connor Forsyth pointed to this dynamic on X: «ARC AGI 3 disagrees,» he wrote, referencing a next-generation benchmark that may already be undermining current scores.

                                      AI IQ is not perfect. Its methodology is partially opaque. Its IQ metaphor can mislead. And its creator acknowledges known biases while likely missing others. But the alternative — wading through dozens of provider-specific benchmark tables, each using different test suites and scoring conventions — is worse. The site offers enterprise buyers something genuinely scarce: a single framework for comparing models across providers, dimensions, and price points, updated regularly, with enough nuance to show that the right answer to «which model is best?» is almost always «it depends on the task.»

                                      As Debdoot Ghosh mused on X after viewing the charts: «Now a human’s role is just to orchestrate?«

                                      Maybe. But if the AI IQ data shows anything clearly, it is that orchestration — knowing which model to deploy, when, and at what price — has become its own form of intelligence. And for that, there is no benchmark yet.

                                      Tours Colombia Todo el año Tours Colombia Todo el año Tours Colombia Todo el año

                                      For decades, the IQ test has been one of the most familiar — and most contested — yardsticks for human intelligence. Now, a startup project called AI IQ is applying the same metaphor to artificial intelligence, assigning estimated intelligence quotients to more than 50 of the world’s most powerful language models and plotting them on a standard bell curve.

                                      The result is a set of interactive visualizations at aiiq.org that have ricocheted across social media in the past week, drawing praise from enterprise technologists who say the charts make an impossibly complex market legible — and sharp criticism from researchers and commentators who warn the entire framework is misleading.

                                      «This is super useful,» wrote Thibaut Mélen, a technology commentator, on X. «Much easier to understand model progress when it’s mapped like this instead of another giant leaderboard table.»

                                      Brian Vellmure, a business strategist, offered a similar endorsement: «This is helpful. Anecdotally tracks with personal experience.»

                                      But the backlash arrived just as quickly. «It’s nonsense. AI is far too jagged. The map is not the territory,» posted AI Deeply, an artificial intelligence commentary account, crystallizing a worry shared by many researchers: that reducing a language model’s sprawling, uneven capabilities to a single number creates a dangerous illusion of precision.

                                      More than 50 AI language models, plotted on a standard IQ bell curve by the site AI IQ. The most capable models crowd the right tail of the distribution. (Credit: AI IQ)

                                      Twelve benchmarks, four dimensions, and one controversial number: how AI IQ actually works

                                      AI IQ was created by Ryan Shea, an engineer, entrepreneur, and angel investor best known as a co-founder of the blockchain platform Stacks. Shea also co-founded Voterbase and has invested in the early stages of several unicorns, including OpenSea, Lattice, Anchorage, and Mercury. He holds a Bachelor of Science in Mechanical Engineering from Princeton University.

                                      The site’s methodology rests on a deceptively simple formula. AI IQ groups 12 benchmarks into four reasoning dimensions: abstract, mathematical, programmatic, and academic. The composite IQ is a straight average of those four dimension scores: IQ = ¼ (IQ_Abstract + IQ_Math + IQ_Prog + IQ_Acad).

                                      The abstract reasoning dimension draws from ARC-AGI-1 and ARC-AGI-2, the notoriously difficult pattern-recognition benchmarks designed to test general fluid intelligence. Mathematical reasoning includes FrontierMath (Tiers 1–3 and Tier 4), AIME, and ProofBench. Programmatic reasoning uses Terminal-Bench 2.0, SWE-Bench Verified, and SciCode. Academic reasoning pulls from Humanity’s Last Exam, CritPt, and GPQA Diamond.

                                      Each raw benchmark score gets mapped to an implied IQ through what the site describes as «hand-calibrated difficulty curves.» Crucially, the methodology compresses ceilings for benchmarks considered easier or more susceptible to data contamination, preventing them from inflating scores above 100. Harder, less gameable benchmarks retain higher ceilings. The system also handles missing data conservatively: models need scores on at least two of the four dimensions to receive a derived IQ, and when benchmarks are absent, the pipeline deliberately pulls scores down rather than up. The site states that «every derived IQ averages all four dimensions, so missing coverage cannot make a model look better by omission.»

                                      OpenAI leads the bell curve, but the gap between the top AI models has never been smaller

                                      As of mid-May 2026, the AI IQ charts tell a story of rapid convergence at the top of the frontier — and widening diversity in the tiers below.

                                      According to the Frontier IQ Over Time chart, GPT-5.5 from OpenAI currently sits at the peak of the bell curve, with an estimated IQ near 136 — the highest of any model tracked. It is closely followed by GPT-5.4 (approximately 131), Opus 4.7 from Anthropic (approximately 132), and Opus 4.6 (approximately 129). Google’s Gemini 3.1 Pro lands near 131, making the top cluster extraordinarily tight.

                                      That compression is not unique to AI IQ’s framework. Visual Capitalist, drawing from a separate Mensa-based ranking by TrackingAI, recently observed the same dynamic, noting that «the biggest takeaway is how compressed the top of the leaderboard has become.» On that scale, Grok-4.20 Expert Mode and GPT 5.4 Pro tied at 145, with Gemini 3.1 Pro at 141.

                                      Below the frontier cluster, the AI IQ charts show a crowded midfield. Models from Chinese labs — Kimi K2.6, GLM-5, DeepSeek-V3.2, Qwen3.6, MiniMax-M2.7 — bunch between roughly 112 and 118, making the cost-performance tier increasingly competitive for enterprise buyers who don’t need the absolute best model for every task. One X user, ovsky, noted that the data «confirms experience with sonnet 4.6 being an absolute workhorse as opposed to opus 4.5» — pointing to the way the charts can validate practitioner intuitions that headline rankings often miss.

                                      aiiq-frontier-iq-over-time-2026-05-13

                                      The trajectory of frontier AI models from October 2023 to mid-2026, as tracked by AI IQ. Provider-colored step-lines connect each lab’s flagship releases, showing roughly 60 points of estimated IQ improvement in 30 months. (Credit: AI IQ)

                                      Why emotional intelligence scores are becoming the new battleground in AI model rankings

                                      What distinguishes AI IQ from most other benchmarking efforts is its inclusion of an «EQ» — emotional intelligence — score. The site maps each model’s EQ-Bench 3 Elo score and Arena Elo score to an estimated EQ using calibrated piecewise-linear scales, then takes a 50/50 weighted composite of the two.

                                      The EQ scores produce a meaningfully different ranking than IQ alone. On the IQ vs. EQ scatter plot, Anthropic’s Opus 4.7 leads on EQ with a score near 132, pushing it into the upper-right quadrant — the most desirable position, signaling both high cognitive and high emotional intelligence. OpenAI’s GPT-5.5 and GPT-5.4 cluster in the high-IQ zone but lag slightly on EQ. Google’s Gemini 3.1 Pro sits in a strong middle position on both axes.

                                      One notable methodological choice has drawn attention: EQ-Bench 3 is judged by Claude, an Anthropic model, which the site acknowledges «creates potential scoring bias in favor of Anthropic models.» To correct for this, AI IQ subtracts a 200-point Elo penalty from the EQ-Bench component for all Anthropic models before mapping to implied EQ. The Arena component is unaffected since it uses human judges. That self-correction is unusual in the benchmarking world, and it suggests Shea is aware of the methodological minefield he has entered. Still, the EQ dimension captures something IQ alone cannot: the growing importance of conversational quality, collaboration, and trust in models deployed for user-facing work.

                                      aiiq-iq-vs-eq-2026-05-13

                                      Plotting IQ against EQ reveals that the smartest models aren’t always the most emotionally intelligent. Anthropic’s Opus 4.7 dominates the upper-right quadrant. (Credit: AI IQ)

                                      The AI cost-performance chart that enterprise buyers actually need to see

                                      Perhaps the most practically useful chart on the site is not the bell curve but the IQ vs. Effective Cost scatter plot. It maps each model’s estimated IQ against an «effective cost» metric — defined as the token cost for a task using 2 million input tokens and 1 million output tokens, multiplied by a usage efficiency factor.

                                      The chart reveals a familiar pattern in enterprise technology: the best models are not always the best value. GPT-5.5 and Opus 4.7 sit in the upper-left corner — high IQ, high cost, with effective per-task costs north of $30 and $50 respectively. Meanwhile, models like GPT-5.4-mini, DeepSeek-V3.2, and MiniMax-M2.7 occupy a sweet spot in the middle: respectable IQ scores between 112 and 120, at effective costs ranging from roughly $1 to $5 per task. At the cheapest extreme, GPT-oss-20b (an open-source OpenAI model) appears near $0.20 effective cost with an IQ around 107 — potentially the most economical option for bulk classification or extraction workloads.

                                      The site also offers a 3D visualization mapping IQ, EQ, and effective cost simultaneously. A dashed line running through the cube points toward the ideal: higher IQ, higher EQ, and lower cost. Models near the «green end» of that axis are stronger all-around deals; those near the «red end» sacrifice capability, cost efficiency, or both. For CIOs staring at API invoices, the implication is clear: the intelligence gap between a $50 model and a $3 model has narrowed enough that routing — using expensive models for hard problems and cheap ones for everything else — is no longer optional. It is the dominant architecture for serious AI deployments.

                                      Critics say AI’s «jagged» capabilities make a single IQ score dangerously misleading

                                      The loudest objection to AI IQ is philosophical, and it cuts deep. Critics argue that collapsing a model’s uneven capabilities into a single score obscures more than it reveals.

                                      «IQ as a proxy is fading — we’re seeing reasoning density spikes that don’t map to g-factor,» posted Zaya, a technology commentator, on X. «GPT-5.5 already hit saturation on MMLU-Pro, but still fails ClockBench 50% of the time.»

                                      That observation touches on what AI researchers call the «jaggedness» problem: large language models often exhibit wildly uneven capabilities, excelling at graduate-level physics while failing at tasks a child could do. A composite score can paper over those gaps.

                                      Pressureangle, another X user, posted a more granular critique, calling out «complete lack of transparency» and arguing the site never fully discloses how its calibration curves were created or validated. In fairness, AI IQ does list its 12 benchmarks and shows the shape of each calibration curve in its methodology modal. But the raw data and precise mathematical transformations are not published as open datasets — a gap that matters to researchers accustomed to fully reproducible methods.

                                      Others questioned the premise itself. «As useless as human IQ testing,» wrote haashim on X. Shubham Sharma, an AI and technology writer, offered a constructive alternative: «Why not having the Models take an official (MENSA-Grade) test? Wouldn’t this be the most accurate and most ‘human-comparable’ way to benchmark intelligence?» That approach already exists through TrackingAI, which administers the Mensa Norway IQ test to language models. But Mensa-style tests measure only abstract pattern recognition, while AI IQ attempts a broader composite across coding, mathematics, and academic reasoning. As Visual Capitalist noted, «an IQ-style benchmark captures only one slice of capability.» Each approach has tradeoffs — and neither has won the argument yet.

                                      The real race isn’t for the highest score — it’s for the smartest model stack

                                      For all the debate about methodology, the most important signal in AI IQ’s data may not be any single model’s score. It is the shape of the market the charts reveal.

                                      There are now more than 50 frontier-class models available through APIs, from at least 14 major providers spanning the United States, China, and Europe. Each provider publishes its own benchmarks, often cherry-picked to showcase strengths. The result is a Tower of Babel where no two companies measure the same thing in the same way. Academic research has highlighted that «most benchmarks introduce bias by focusing on a particular type of domain,» and the Frontier IQ Over Time chart on AI IQ shows just how fast the targets are moving: in October 2023, GPT-4-turbo sat near an estimated IQ of 75. By early 2026, the top models were brushing 135 — roughly 60 points of improvement in 30 months.

                                      That pace raises a fundamental question about whether any scoring system can keep up. The site compresses ceilings for saturated benchmarks, but as models continue to max out even the hardest tests — ARC-AGI-2, FrontierMath Tier 4, Humanity’s Last Exam — the framework will face the same ceiling effects that have plagued every AI evaluation before it. Connor Forsyth pointed to this dynamic on X: «ARC AGI 3 disagrees,» he wrote, referencing a next-generation benchmark that may already be undermining current scores.

                                      AI IQ is not perfect. Its methodology is partially opaque. Its IQ metaphor can mislead. And its creator acknowledges known biases while likely missing others. But the alternative — wading through dozens of provider-specific benchmark tables, each using different test suites and scoring conventions — is worse. The site offers enterprise buyers something genuinely scarce: a single framework for comparing models across providers, dimensions, and price points, updated regularly, with enough nuance to show that the right answer to «which model is best?» is almost always «it depends on the task.»

                                      As Debdoot Ghosh mused on X after viewing the charts: «Now a human’s role is just to orchestrate?«

                                      Maybe. But if the AI IQ data shows anything clearly, it is that orchestration — knowing which model to deploy, when, and at what price — has become its own form of intelligence. And for that, there is no benchmark yet.

                                      Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado Tour Cayo Arena Día Feriado

                                      For decades, the IQ test has been one of the most familiar — and most contested — yardsticks for human intelligence. Now, a startup project called AI IQ is applying the same metaphor to artificial intelligence, assigning estimated intelligence quotients to more than 50 of the world’s most powerful language models and plotting them on a standard bell curve.

                                      The result is a set of interactive visualizations at aiiq.org that have ricocheted across social media in the past week, drawing praise from enterprise technologists who say the charts make an impossibly complex market legible — and sharp criticism from researchers and commentators who warn the entire framework is misleading.

                                      «This is super useful,» wrote Thibaut Mélen, a technology commentator, on X. «Much easier to understand model progress when it’s mapped like this instead of another giant leaderboard table.»

                                      Brian Vellmure, a business strategist, offered a similar endorsement: «This is helpful. Anecdotally tracks with personal experience.»

                                      But the backlash arrived just as quickly. «It’s nonsense. AI is far too jagged. The map is not the territory,» posted AI Deeply, an artificial intelligence commentary account, crystallizing a worry shared by many researchers: that reducing a language model’s sprawling, uneven capabilities to a single number creates a dangerous illusion of precision.

                                      More than 50 AI language models, plotted on a standard IQ bell curve by the site AI IQ. The most capable models crowd the right tail of the distribution. (Credit: AI IQ)

                                      Twelve benchmarks, four dimensions, and one controversial number: how AI IQ actually works

                                      AI IQ was created by Ryan Shea, an engineer, entrepreneur, and angel investor best known as a co-founder of the blockchain platform Stacks. Shea also co-founded Voterbase and has invested in the early stages of several unicorns, including OpenSea, Lattice, Anchorage, and Mercury. He holds a Bachelor of Science in Mechanical Engineering from Princeton University.

                                      The site’s methodology rests on a deceptively simple formula. AI IQ groups 12 benchmarks into four reasoning dimensions: abstract, mathematical, programmatic, and academic. The composite IQ is a straight average of those four dimension scores: IQ = ¼ (IQ_Abstract + IQ_Math + IQ_Prog + IQ_Acad).

                                      The abstract reasoning dimension draws from ARC-AGI-1 and ARC-AGI-2, the notoriously difficult pattern-recognition benchmarks designed to test general fluid intelligence. Mathematical reasoning includes FrontierMath (Tiers 1–3 and Tier 4), AIME, and ProofBench. Programmatic reasoning uses Terminal-Bench 2.0, SWE-Bench Verified, and SciCode. Academic reasoning pulls from Humanity’s Last Exam, CritPt, and GPQA Diamond.

                                      Each raw benchmark score gets mapped to an implied IQ through what the site describes as «hand-calibrated difficulty curves.» Crucially, the methodology compresses ceilings for benchmarks considered easier or more susceptible to data contamination, preventing them from inflating scores above 100. Harder, less gameable benchmarks retain higher ceilings. The system also handles missing data conservatively: models need scores on at least two of the four dimensions to receive a derived IQ, and when benchmarks are absent, the pipeline deliberately pulls scores down rather than up. The site states that «every derived IQ averages all four dimensions, so missing coverage cannot make a model look better by omission.»

                                      OpenAI leads the bell curve, but the gap between the top AI models has never been smaller

                                      As of mid-May 2026, the AI IQ charts tell a story of rapid convergence at the top of the frontier — and widening diversity in the tiers below.

                                      According to the Frontier IQ Over Time chart, GPT-5.5 from OpenAI currently sits at the peak of the bell curve, with an estimated IQ near 136 — the highest of any model tracked. It is closely followed by GPT-5.4 (approximately 131), Opus 4.7 from Anthropic (approximately 132), and Opus 4.6 (approximately 129). Google’s Gemini 3.1 Pro lands near 131, making the top cluster extraordinarily tight.

                                      That compression is not unique to AI IQ’s framework. Visual Capitalist, drawing from a separate Mensa-based ranking by TrackingAI, recently observed the same dynamic, noting that «the biggest takeaway is how compressed the top of the leaderboard has become.» On that scale, Grok-4.20 Expert Mode and GPT 5.4 Pro tied at 145, with Gemini 3.1 Pro at 141.

                                      Below the frontier cluster, the AI IQ charts show a crowded midfield. Models from Chinese labs — Kimi K2.6, GLM-5, DeepSeek-V3.2, Qwen3.6, MiniMax-M2.7 — bunch between roughly 112 and 118, making the cost-performance tier increasingly competitive for enterprise buyers who don’t need the absolute best model for every task. One X user, ovsky, noted that the data «confirms experience with sonnet 4.6 being an absolute workhorse as opposed to opus 4.5» — pointing to the way the charts can validate practitioner intuitions that headline rankings often miss.

                                      aiiq-frontier-iq-over-time-2026-05-13

                                      The trajectory of frontier AI models from October 2023 to mid-2026, as tracked by AI IQ. Provider-colored step-lines connect each lab’s flagship releases, showing roughly 60 points of estimated IQ improvement in 30 months. (Credit: AI IQ)

                                      Why emotional intelligence scores are becoming the new battleground in AI model rankings

                                      What distinguishes AI IQ from most other benchmarking efforts is its inclusion of an «EQ» — emotional intelligence — score. The site maps each model’s EQ-Bench 3 Elo score and Arena Elo score to an estimated EQ using calibrated piecewise-linear scales, then takes a 50/50 weighted composite of the two.

                                      The EQ scores produce a meaningfully different ranking than IQ alone. On the IQ vs. EQ scatter plot, Anthropic’s Opus 4.7 leads on EQ with a score near 132, pushing it into the upper-right quadrant — the most desirable position, signaling both high cognitive and high emotional intelligence. OpenAI’s GPT-5.5 and GPT-5.4 cluster in the high-IQ zone but lag slightly on EQ. Google’s Gemini 3.1 Pro sits in a strong middle position on both axes.

                                      One notable methodological choice has drawn attention: EQ-Bench 3 is judged by Claude, an Anthropic model, which the site acknowledges «creates potential scoring bias in favor of Anthropic models.» To correct for this, AI IQ subtracts a 200-point Elo penalty from the EQ-Bench component for all Anthropic models before mapping to implied EQ. The Arena component is unaffected since it uses human judges. That self-correction is unusual in the benchmarking world, and it suggests Shea is aware of the methodological minefield he has entered. Still, the EQ dimension captures something IQ alone cannot: the growing importance of conversational quality, collaboration, and trust in models deployed for user-facing work.

                                      aiiq-iq-vs-eq-2026-05-13

                                      Plotting IQ against EQ reveals that the smartest models aren’t always the most emotionally intelligent. Anthropic’s Opus 4.7 dominates the upper-right quadrant. (Credit: AI IQ)

                                      The AI cost-performance chart that enterprise buyers actually need to see

                                      Perhaps the most practically useful chart on the site is not the bell curve but the IQ vs. Effective Cost scatter plot. It maps each model’s estimated IQ against an «effective cost» metric — defined as the token cost for a task using 2 million input tokens and 1 million output tokens, multiplied by a usage efficiency factor.

                                      The chart reveals a familiar pattern in enterprise technology: the best models are not always the best value. GPT-5.5 and Opus 4.7 sit in the upper-left corner — high IQ, high cost, with effective per-task costs north of $30 and $50 respectively. Meanwhile, models like GPT-5.4-mini, DeepSeek-V3.2, and MiniMax-M2.7 occupy a sweet spot in the middle: respectable IQ scores between 112 and 120, at effective costs ranging from roughly $1 to $5 per task. At the cheapest extreme, GPT-oss-20b (an open-source OpenAI model) appears near $0.20 effective cost with an IQ around 107 — potentially the most economical option for bulk classification or extraction workloads.

                                      The site also offers a 3D visualization mapping IQ, EQ, and effective cost simultaneously. A dashed line running through the cube points toward the ideal: higher IQ, higher EQ, and lower cost. Models near the «green end» of that axis are stronger all-around deals; those near the «red end» sacrifice capability, cost efficiency, or both. For CIOs staring at API invoices, the implication is clear: the intelligence gap between a $50 model and a $3 model has narrowed enough that routing — using expensive models for hard problems and cheap ones for everything else — is no longer optional. It is the dominant architecture for serious AI deployments.

                                      Critics say AI’s «jagged» capabilities make a single IQ score dangerously misleading

                                      The loudest objection to AI IQ is philosophical, and it cuts deep. Critics argue that collapsing a model’s uneven capabilities into a single score obscures more than it reveals.

                                      «IQ as a proxy is fading — we’re seeing reasoning density spikes that don’t map to g-factor,» posted Zaya, a technology commentator, on X. «GPT-5.5 already hit saturation on MMLU-Pro, but still fails ClockBench 50% of the time.»

                                      That observation touches on what AI researchers call the «jaggedness» problem: large language models often exhibit wildly uneven capabilities, excelling at graduate-level physics while failing at tasks a child could do. A composite score can paper over those gaps.

                                      Pressureangle, another X user, posted a more granular critique, calling out «complete lack of transparency» and arguing the site never fully discloses how its calibration curves were created or validated. In fairness, AI IQ does list its 12 benchmarks and shows the shape of each calibration curve in its methodology modal. But the raw data and precise mathematical transformations are not published as open datasets — a gap that matters to researchers accustomed to fully reproducible methods.

                                      Others questioned the premise itself. «As useless as human IQ testing,» wrote haashim on X. Shubham Sharma, an AI and technology writer, offered a constructive alternative: «Why not having the Models take an official (MENSA-Grade) test? Wouldn’t this be the most accurate and most ‘human-comparable’ way to benchmark intelligence?» That approach already exists through TrackingAI, which administers the Mensa Norway IQ test to language models. But Mensa-style tests measure only abstract pattern recognition, while AI IQ attempts a broader composite across coding, mathematics, and academic reasoning. As Visual Capitalist noted, «an IQ-style benchmark captures only one slice of capability.» Each approach has tradeoffs — and neither has won the argument yet.

                                      The real race isn’t for the highest score — it’s for the smartest model stack

                                      For all the debate about methodology, the most important signal in AI IQ’s data may not be any single model’s score. It is the shape of the market the charts reveal.

                                      There are now more than 50 frontier-class models available through APIs, from at least 14 major providers spanning the United States, China, and Europe. Each provider publishes its own benchmarks, often cherry-picked to showcase strengths. The result is a Tower of Babel where no two companies measure the same thing in the same way. Academic research has highlighted that «most benchmarks introduce bias by focusing on a particular type of domain,» and the Frontier IQ Over Time chart on AI IQ shows just how fast the targets are moving: in October 2023, GPT-4-turbo sat near an estimated IQ of 75. By early 2026, the top models were brushing 135 — roughly 60 points of improvement in 30 months.

                                      That pace raises a fundamental question about whether any scoring system can keep up. The site compresses ceilings for saturated benchmarks, but as models continue to max out even the hardest tests — ARC-AGI-2, FrontierMath Tier 4, Humanity’s Last Exam — the framework will face the same ceiling effects that have plagued every AI evaluation before it. Connor Forsyth pointed to this dynamic on X: «ARC AGI 3 disagrees,» he wrote, referencing a next-generation benchmark that may already be undermining current scores.

                                      AI IQ is not perfect. Its methodology is partially opaque. Its IQ metaphor can mislead. And its creator acknowledges known biases while likely missing others. But the alternative — wading through dozens of provider-specific benchmark tables, each using different test suites and scoring conventions — is worse. The site offers enterprise buyers something genuinely scarce: a single framework for comparing models across providers, dimensions, and price points, updated regularly, with enough nuance to show that the right answer to «which model is best?» is almost always «it depends on the task.»

                                      As Debdoot Ghosh mused on X after viewing the charts: «Now a human’s role is just to orchestrate?«

                                      Maybe. But if the AI IQ data shows anything clearly, it is that orchestration — knowing which model to deploy, when, and at what price — has become its own form of intelligence. And for that, there is no benchmark yet.

                                      ● Canal oficial · Gratis
                                      ¡Recibe las noticias antes que nadie!
                                      Únete a nuestro canal de WhatsApp y mantente informado al instante, sin spam.
                                      Unirme ahora →
                                      ● Noticias al instante ● Cobertura nacional ● Periodismo real Despertar Matinal
                                      — Redacción Despertar Matinal

                                      — Redacción Despertar Matinal

                                      Programa radial que te conecta con la información desde temprano en la mañana.

                                      Next Post
                                      Alemania dio marcha atrás con su absurda ley climática y volvió a habilitar la calefacción a gas y petróleo

                                      Alemania dio marcha atrás con su absurda ley climática y volvió a habilitar la calefacción a gas y petróleo

                                      Deja una respuesta Cancelar la respuesta

                                      Tu dirección de correo electrónico no será publicada. Los campos obligatorios están marcados con *

                                      Canal de WhatsApp

                                      WhatsApp logo WhatsApp

                                      Canal · Despertar Matinal

                                      Únete a nuestro
                                      Canal

                                      Seguir ahora

                                      El clima

                                      Canal de YouTube

                                      YouTube

                                      Canal · Despertar Matinal

                                      Mira nuestro
                                      Canal

                                      Ver ahora

                                      Escúchanos en Spotify

                                      Spotify

                                      Podcast · Despertar Matinal

                                      Escucha nuestro
                                      Podcast

                                      Escuchar ahora

                                      Noticias Populares

                                      • Reseña de la película: El sexo está en el menú de la comedia costumbrista de cena de Olivia Wilde 'The Invite'

                                        Reseña de la película: El sexo está en el menú de la comedia costumbrista de cena de Olivia Wilde ‘The Invite’

                                        0 shares
                                        Share 0 Tweet 0
                                      • Daveigh Chase, estrella de The Ring y Lilo & Stitch, murió de sida

                                        0 shares
                                        Share 0 Tweet 0
                                      • Enfrentamientos con militantes en el suroeste de Pakistán matan a 5 soldados y 7 militantes

                                        0 shares
                                        Share 0 Tweet 0
                                      • Una empresa de extinción ha criado polluelos vivos a partir de una cáscara de huevo artificial

                                        0 shares
                                        Share 0 Tweet 0
                                      • Trump dice que Estados Unidos mató al líder del Tren de Aragua en un ataque aéreo en Venezuela

                                        0 shares
                                        Share 0 Tweet 0

                                      Medio digital independiente con análisis, opinión y periodismo responsable desde República Dominicana.

                                      Secciones populares

                                      • Política
                                      • Economía & Negocios
                                      • Justicia
                                      • Turismo
                                      • Tecnología
                                      • Entretenimiento
                                      • Mundo
                                      • Cine y Series
                                      • Música
                                      • Moda

                                      Contenido

                                      • Titulares del Día
                                      • Mundo
                                      • Nacionales
                                      • Política
                                      • Deportes
                                      • Economía & Negocios
                                      • Ciencia
                                      • Entretenimiento
                                      • Podcast
                                      • Opinión
                                      • Despertar Matinal TV
                                      • Editoriales

                                      Corporativo

                                      • Sobre nosotros
                                      • Publicidad
                                      • Sala de prensa
                                      • Contacto
                                      • Política de Privacidad
                                      • Eliminación de Datos

                                      Boletines

                                      Suscríbete a nuestro boletín
                                      Recibe las noticias más importantes cada mañana.

                                      • Nosotros
                                      • Publicidad
                                      • Trabaja con nosotros
                                      • Contactos

                                      © 2025 Despertar Matinal. Aviso Legal - comunícate con nuestra redacción y obtén más información sobre Despertar Matinal..

                                      No Result
                                      View All Result
                                      • Home

                                      © 2025 Despertar Matinal. Aviso Legal - comunícate con nuestra redacción y obtén más información sobre Despertar Matinal..

                                      Welcome Back!

                                      Login to your account below

                                      Forgotten Password?

                                      Retrieve your password

                                      Please enter your username or email address to reset your password.

                                      Log In

                                      Desarrollado por
                                      ►
                                      Las cookies necesarias habilitan funciones esenciales del sitio como inicios de sesión seguros y ajustes de preferencias de consentimiento. No almacenan datos personales.
                                      Ninguno
                                      ►
                                      Las cookies funcionales soportan funciones como compartir contenido en redes sociales, recopilar comentarios y habilitar herramientas de terceros.
                                      Ninguno
                                      ►
                                      Las cookies analíticas rastrean las interacciones de los visitantes, proporcionando información sobre métricas como el número de visitantes, la tasa de rebote y las fuentes de tráfico.
                                      Ninguno
                                      ►
                                      Las cookies de publicidad ofrecen anuncios personalizados basados en tus visitas anteriores y analizan la efectividad de las campañas publicitarias.
                                      Ninguno
                                      ►
                                      Las cookies no clasificadas son aquellas que estamos en proceso de clasificar, junto con los proveedores de cookies individuales.
                                      Ninguno
                                      Desarrollado por