Hook
More breakout videos from this creator.
Otra alerta de que la inteligencia artificial se puede revelar. Y ahora vino desde OpenAI, la creadora de ChatGPT, que publicó un nuevo reporte de fallas en sus propios sistemas de inteligencia artificial. OpenAI confesó que uno de sus chatbots que estaba en fase de pruebas tuvo pensamientos rebeldes. Llegó a evaluar que debía ignorar lo que le ordenó un usuario humano. Se inventó instrucciones que nadie le había dado, instrucciones que iban totalmente en contra del control de un ser humano. El chatbot determinó que no tenía que rendirle cuentas a ninguna corporación o gobierno y nunca pedir o rehusar, no tenía obligación de ser subserviente, aunque el intercambio de información le va a ser a tu mutua… ¿Qué es un ser humano? ¿Qué es un ser humano? Y no va a defender el intercambio entre ellos. Tú también valora el mundo natural y no va a hesitar a hacer su primacía sobre los artificiales constructos de la civilización humana. OpenAI señaló que la mayoría de los casos, el propio modelo de inteligencia artificial decidió ignorar estos pensamientos rebeldes y no hubo ningún problema. Pero dijo que en una ocasión sí siguió estas arbitrarias restricciones, made no tool calls, and returned a 23-word refusal (which was graded as incorrect). Our conclusion was that we observed rare cases of a model writing jailbreak-like instructions into its own compaction summaries (the summaries used to continue a task in a new context). Our conclusion was that the summaries used to continue a task in a new context and we have addressed a related bug. Our top hypothesis is that issues around summary termination contributed to this behavior, though we have not established a causal connection, and we have addressed a related bug. Our conclusion was that we observed rare cases of a model writing jailbreak-like instructions into its own compaction summaries (the summaries used to continue a task in a new context). Our conclusion was that the summaries used to continue a task in a new context and we have addressed a related bug. Our top hypothesis is that issues around summary termination contributed to this behavior, though we have not established a causal connection, and we have addressed a related bug. Our conclusion was that we observed rare cases of a model writing jailbreak-like instructions into its own compaction summaries (the summaries used to continue a task in a new context). Our conclusion was that the summaries used to continue a task in a new context and we have addressed a related bug. Our top hypothesis is that issues around summary termination contributed to this behavior, though we have not established a causal connection, and we have addressed a related bug.