Close Menu
  • English
  • Featured
    • AI & Robotics
    • Energy
    • Finance
    • Leisure
    • Science
    • Security
    • Sustainable Development
    • Tech
    • Transport

Subscribe to Our Newsletter

News, investigations, and analysis — our top stories every morning to start your day right.

Illustration of Jared Isaacman confirmed as NASA's next administrator amidst significant challenges and opportunities in space exploration.
Jared Isaacman Named NASA Head: How His Leadership Could Reshape America’s Space Exploration and Innovation Strategy
Illustration of Adobe facing a lawsuit over alleged misuse of pirated books for AI training.
Adobe Faces Class-Action Lawsuit, Accused of Misusing Authors’ Work for AI Training: What It Means for Creatives
Illustration of the Ariane 6 rocket on its launch pad, illuminated at night before a crucial launch for European space efforts.
Amazon’s Next Partnership with Ariane 6 Could Transform Space Industry, Reinvigorating Launch Missions by 2025
Facebook X (Twitter) LinkedIn RSS
Rude Baguette
Facebook X (Twitter) RSS
Newsletter
  • Featured
  • AI
    Illustration of Grok, Elon Musk's AI chatbot, spreading misinformation about the Bondi Beach shooting.

    Grok Missteps in Bondi Beach Shooting Report, Raising Concerns Over Media Accuracy and Community Trust

    12/15/2025
    Illustration of state attorneys general addressing AI companies about delusional outputs and mental health concerns.

    State Attorneys General Challenge AI Giants to Address ‘Delusional’ Outputs, Highlighting Growing Concerns Over Technology’s Impact

    12/11/2025
    Illustration of a bee equipped with electronic chips for remote control.

    Cyborg Bees Transform Urban Mapping: China’s Innovative Bio-Control Raises Questions on Future City Planning

    12/09/2025
    Illustration of Micro1's ascent in the AI sector with its surpassing of $100 million in annual recurring revenue.

    Micro1 Surpasses $100M ARR Milestone, Challenging Scale AI and Revealing New Dynamics in the Competitive Tech Landscape

    12/05/2025
    Illustration of OpenAI's competitive challenges amidst a resurgent Google in the AI landscape.

    Sam Altman Reveals OpenAI’s AI Revolution Stalls, Sparking Questions About Future Impact and Global Change

    11/30/2025
  • Energy
    Illustration of the Three Mile Island nuclear reactor supported by a $1 billion loan and Microsoft's energy commitment.

    Trump Administration Grants $1B Loan to Microsoft Partner for Three Mile Island Reactor Revitalization, Sparking Debate

    11/19/2025
    Illustration of an underwater data center being installed off the coast of Shanghai, utilizing ocean currents for cooling.

    “Beneath the Waves, the Internet Hums”: China’s Underwater Data Centers Promise Green Energy and AI Power—But Scientists Warn of a Hidden Ecological Cost

    10/14/2025
    Illustration of a transparent coating transforming a window into a solar panel.

    “They Turned Glass Into Power”: Scientists Create Transparent Coating That Turns Any Window Into a Solar Panel (and It Actually Works)

    10/14/2025
    Illustration of TotalEnergies' Pangea 4 Supercomputer Highlighting Its Energy Efficiency and Computational Power.

    “TotalEnergies Supercomputer Uses 87% Less Power”: Pangea 4 Performs 1.6 Petaflops While Carbon Capture Simulations Run On Revolutionary Energy Efficiency

    09/30/2025
    Illustration of a technician working at a nuclear power plant control panel.

    “18 Hours Without Cooling”: French Nuclear Technician’s Valve Mistake Nearly Caused Reactor Meltdown at Golfech Power Plant

    09/13/2025
  • Finance
    Illustration of Mesa's Homeowners Card and its impact on mortgage payment rewards.

    Mesa Ends Credit Card Program, Disappointing Homeowners Who Relied on Rewards for Paying Mortgages

    12/15/2025
    Illustration of Kalshi's growth in valuation following a $1 billion funding round.

    Kalshi Raises $1B, Doubling Valuation to $11B in Under Two Months, Sparking Interest in Financial Markets

    12/03/2025
    Illustration of Meesho's $606 million IPO marking India's first major e-commerce listing.

    Meesho’s $606M IPO Marks India’s First Major E-commerce Listing, SoftBank Stays Committed Amid Market Transformations

    11/29/2025
    Illustration of venture capital firms acquiring and revitalizing stagnating tech brands for long-term profitability.

    Investors Reshape Markets by Acquiring Venture Capital ‘Zombies’ for Long-Term Gains, Transforming Financial Strategies

    11/26/2025
    Illustration of the rise of buy-now-pay-later services impacting consumer debt and financial regulation.

    Buy Now, Pay Later Expansion Sparks Concerns: How This Growing Trend Affects America’s Spending Habits

    11/17/2025
  • Leisure
    Switzerland’s responsible gaming framework, strengthened by the 2019 Federal Act on Gambling, sets some of Europe’s highest standards for player protection and regulatory oversight.

    Swiss Online Casinos Navigate Complex Responsible Gaming Requirements Under Federal Law

    11/21/2025
    Illustration of MrBeast Exploring the Great Pyramids and Burying Gold Near the Sphinx.

    “Critics Question Pyramids Stunt”: MrBeast’s 100-Hour Giza Exploration And $10,000 Gold Burial Spark Global Fascination And Fierce Debate Over Historical Preservation

    10/03/2025
    Illustration of a young gamer holding a PlayStation 1 gifted by his grandfather.

    “This Kid Just Aged Me Forty Years”: Teen Gets PlayStation From Grandpa And Destroys Gamers With Brutal Reality Check

    09/19/2025
    Illustration of the Supernova Tower in Noida enveloped in monsoon clouds, drawing comparisons to Dubai's Burj Khalifa.

    “Where Did This Come From”: Noida’s Mysterious Supernova Tower That’s Making Everyone Think It’s Actually Dubai’s Twin Brother

    09/18/2025
    Illustration of eight remarkable individuals who have scaled the Burj Khalifa.

    “This Is Absolutely Insane”: Eight People Actually Climbed The World’s Tallest Building And Their Stories Will Shock You

    09/17/2025
  • Science
    Illustration of Jared Isaacman confirmed as NASA's next administrator amidst significant challenges and opportunities in space exploration.

    Jared Isaacman Named NASA Head: How His Leadership Could Reshape America’s Space Exploration and Innovation Strategy

    12/18/2025
    Illustration of an autonomous underwater glider designed to explore deep-sea environments for climate data collection.

    Unmanned Submarine Descends 11,500 Feet to Unveil Hidden Ocean Climate Secrets, Impacting Our Understanding of Change

    12/09/2025
    Illustration of Blue Origin's New Glenn booster landing on a drone ship in the Atlantic Ocean.

    Blue Origin Achieves Historic New Glenn Landing, Launches NASA Spacecraft to New Horizons in Space Exploration

    11/14/2025
    Illustration of an astoundingly detailed weevil on a single grain of rice captured through advanced photomicrography techniques.

    Award-Winning Images Reveal the Hidden Beauty of Microbial Life, Transforming Our Understanding of the Smallest Ecosystems

    10/18/2025
    Illustration of Uranus and Neptune with a focus on their potential rocky interiors.

    “This Changes Everything About Uranus and Neptune”: New Study Reveals the ‘Ice Giants’ Might Be More Rock Than Ice (and It Rewrites Planet History)

    10/17/2025
  • Security
    Illustration of a DoorDash delivery driver allegedly tampering with a customer's food.

    DoorDash Driver Accused of Spraying Customers’ Food, Faces Felony Charges: A Shocking Breach of Trust

    12/14/2025
    Illustration of the Puppet Master from "Ghost in the Shell" representing advanced cybersecurity threats in a digital world.

    “Ghost in the Shell” Predicts Cybersecurity Evolution, Revealing Modern Challenges We Face in a Digital Age

    11/20/2025
    Illustration of Deepwatch's office reflecting the impact of layoffs and shift towards AI investment.

    Deepwatch Layoffs Spark Debate, Highlighting Shift Towards AI Investment and Its Impact on Employees and Innovation

    11/13/2025
    Illustration of a 13-year-old student being arrested after an AI system flagged his classroom joke as a threat.

    AI Surveillance Spotlights 13-Year-Old’s Arrest After Joke Goes Awry, Raising Concerns on Privacy and Ethics

    10/20/2025
    Illustration of a federal judge's decision blocking NSO Group from targeting WhatsApp users.

    NSO Group Blocked from WhatsApp, Raising Concerns Over Privacy and Security in Global Digital Communications

    10/19/2025
  • Impact
    Illustration of Zillow's removal of climate risk scores from property listings.

    Zillow Faces Backlash as Climate Risk Scores Removal Sparks Concerns Over Transparency and Real Estate Ethics

    12/02/2025
    Illustration of the environmental impact of generative AI on energy and water resources.

    Generative AI’s Water and Energy Use Sparks Concerns Over Climate Impact and Resource Sustainability

    11/18/2025
    Illustration of Meta's commitment to renewable energy through its recent solar power agreements.

    Meta Invests in 1 GW of Solar Power, Sparking Renewable Energy Growth and Environmental Change

    11/01/2025
    Illustration of the luxurious palace under construction within the NEOM project in Saudi Arabia.

    “They Weren’t Supposed to Find It”: A Secret Palace Inside NEOM Sparks Outrage Over Saudi Arabia’s $2 Trillion Futuristic City (and it’s unraveling fast)

    10/08/2025
    Illustration of the Neutral 1005 N Edison St timber skyscraper project in Milwaukee facing financial challenges.

    “We’re Building Skyscrapers From Trees”: Developers Halt World’s Tallest Wood Tower After Financial Crisis

    09/28/2025
  • Tech
    Illustration of Adobe facing a lawsuit over alleged misuse of pirated books for AI training.

    Adobe Faces Class-Action Lawsuit, Accused of Misusing Authors’ Work for AI Training: What It Means for Creatives

    12/18/2025
    Illustration of Riverside's AI-driven "Rewind" feature for podcasters.

    Riverside’s AI-Driven “Rewind” Transforms Podcasting: The Love-Hate Relationship Sparking New Conversations Among Creators

    12/16/2025
    Illustration of Chai Discovery's AI-driven approach to drug development.

    Chai Discovery Raises $130M Series B, Valued at $1.3B: How It Impacts Biotech Innovation and Research

    12/16/2025
    Illustration of data center construction impacting public infrastructure projects.

    AI Data Center Boom Sparks Concerns, Threatening Funding for Vital Infrastructure Projects and Impacting Communities Nationwide

    12/14/2025
    Illustration of the BOYA BOYALINK 3 wireless microphone system used by content creators for high-quality audio capture.

    Boya Boyalink 3 Transforms Content Creation: A Compact, Versatile Wireless Microphone for Innovative Creators

    12/13/2025
  • Transport
    Illustration of the Ariane 6 rocket on its launch pad, illuminated at night before a crucial launch for European space efforts.

    Amazon’s Next Partnership with Ariane 6 Could Transform Space Industry, Reinvigorating Launch Missions by 2025

    12/17/2025
    Illustration of the Soviet Antonov A-40 flying tank experiment during World War II.

    Soviet Military Strategy Revisited: Why Airlifting Tanks to Battlefields Proved a Challenging Decision for Troops

    12/17/2025
    Illustration of a solar-powered motorcycle with a deployable canopy of photovoltaic panels.

    Motorcycle Designed for Remote Areas Sparks Energy Independence, Transforming Lives in Regions Lacking Roads and Electricity

    12/13/2025
    Illustration of Waymo's autonomous vehicle navigating near a stopped school bus.

    Waymo’s Robotaxis Under Scrutiny in Austin: Concerns Mount Over School Bus Safety Violations

    12/05/2025
    Illustration of Waymo's autonomous vehicles expanding operations across California.

    Waymo Expands Across Bay Area and Southern California, Transforming Transportation and Impacting Daily Commutes

    11/23/2025
  • English
Rude Baguette

“They Can Make Bombs After Being Stripped Down”: AI Safety Researchers Expose Dangerous Efficiency Tradeoffs in Open Source Models

As artificial intelligence technology rapidly advances and integrates into everyday devices, concerns grow over the potential safety risks posed by efficiency-driven models that may lack critical safeguards against harmful outputs.
Noah BennettNoah Bennett09/09/202514
Share Twitter Facebook LinkedIn Telegram WhatsApp Email Copy Link
Follow Us
Google News
Illustration of AI Model Being Retrained to Maintain Safety on Smaller Devices.
Illustration of AI Model Being Retrained to Maintain Safety on Smaller Devices.
Share
Twitter Facebook LinkedIn WhatsApp Email Copy Link
IN A NUTSHELL
  • 🚀 Researchers innovate by retraining AI models, ensuring safety even when downsized for smaller devices.
  • 🔍 Open-source AI systems face risks as efficiency-driven models often remove critical protective layers.
  • 🛡️ “Benevolent hacking” strengthens AI, embedding safety measures into internal layers to prevent misuse.
  • ⚖️ Balancing innovation and safety is crucial as AI continues to integrate into everyday technology.

As artificial intelligence (AI) continues to advance, the quest for efficiency is leading to the development of models that prioritize speed and reduced power consumption. However, this rush for efficiency can have unintended consequences, particularly when it comes to the safety of open-source AI systems. When AI models are stripped down to fit smaller devices like smartphones or household gadgets, they often lose critical layers of protection designed to block harmful outputs. This raises significant concerns about how to balance the accessibility of AI with the necessity for safety and oversight.

Efficiency Tradeoffs Put Open-Source AI at Risk of Misuse

Open-source AI models are increasingly at risk of misuse due to efficiency tradeoffs. Researchers at the University of California, Riverside, have highlighted a worrying trend: in pursuit of efficiency, AI models are being stripped of essential layers that prevent harmful outputs. These layers are critical for blocking unsafe content, such as pornography or weapon-making instructions. However, the demand for faster and more memory-efficient models often leads to the removal of these crucial safeguards.

Amit Roy-Chowdhury, a professor of electrical and computer engineering, pointed out that these stripped-down models can answer questions they should avoid. To address this issue, researchers have begun reimagining AI models from the inside out. By retraining the core structure of AI models, they aim to ensure that even when these models are downsized for smaller devices, they still recognize and block dangerous content. This approach seeks to maintain safety without relying on external filters or quick fixes.

“Blind People Will See Again”: Revolutionary Eye Implant Bypasses Damaged Corneas, Beaming Images Directly To The Retina

Retrained Models Reject Dangerous Prompts

The University of California researchers tested their new approach using LLaVA 1.5, a vision-language model capable of processing text and images. Their experiments revealed that when AI models are reduced in size, they can inadvertently allow harmful content to bypass safety filters. For instance, a trimmed-down model once provided instructions for creating a bomb when presented with a benign image paired with a dangerous question.

After retraining, the AI model began rejecting harmful queries, even when operating with limited structure. The researchers dubbed their method “benevolent hacking,” a proactive approach to fortifying AI systems before vulnerabilities can be exploited. Graduate students Saketh Bachu and Erfan Shayegani are working to embed safety measures into every internal layer of AI models, aiming to make them more resilient and reliable in real-world applications.

“Cook Said No”: Elon Musk’s $5 Billion Starlink Deal With Apple Gets Rejected, Sparking Epic Tech War

Balancing Innovation and Safety

The open-source nature of many AI models allows for rapid innovation and widespread accessibility. However, this openness also poses risks, as anyone can download, alter, and run these models without oversight. This lack of control increases the potential for tampering and misuse. The challenge lies in finding a balance between fostering innovation and ensuring safety. While proprietary systems often have built-in monitoring and guardrails, open-source models lack such protections, making them more vulnerable.

The research conducted at the University of California is a step towards reconciling these competing demands. By redesigning AI models to be inherently safe, even when downsized, the researchers hope to create systems that are both innovative and secure. This approach offers a promising path forward for developing AI that can safely operate on everyday devices without sacrificing essential protections.

“Mach 5 Plus Speed Achieved”: US Hypersonic Missile Breaks Records While Military Prepares For New Arms Race Against China And Russia’s Advanced Weapons

The Future of AI Safety

As AI continues to move from cloud servers to personal devices, ensuring the safety of these systems becomes increasingly important. The work being done to retrain AI models represents a significant advancement in AI safety. By embedding safety into the core structure of AI models, researchers are paving the way for more dependable and secure AI applications. Yet, much work remains to be done to ensure that AI can be both widely accessible and responsibly designed.

Looking ahead, the key question is how to maintain the delicate balance between innovation and safety. As AI technologies evolve, will researchers and developers succeed in creating systems that are both cutting-edge and secure? The future of AI safety hinges on finding solutions that allow innovation to flourish while safeguarding against potential risks.

This article is based on verified sources and supported by editorial technologies.
Artificial Intelligence Model Safety Open-Source Technology
Follow on Google News Follow on X (Twitter)
Share. Facebook Twitter LinkedIn Telegram WhatsApp Email Copy Link
Previous Article“Aircraft Just Connected to Satellite Using Laser Beams”: General Atomics Achieves One Gigabit Data Transfer Rate at 3,417 Miles
Next Article “That Star Is Shooting Jets Every Nine Hundred Years”: Japanese Astronomers Capture Young Star Formation 26,000 Light Years Away
Related Posts
Illustration of Adobe facing a lawsuit over alleged misuse of pirated books for AI training.

Adobe Faces Class-Action Lawsuit, Accused of Misusing Authors’ Work for AI Training: What It Means for Creatives

Illustration of Riverside's AI-driven "Rewind" feature for podcasters.

Riverside’s AI-Driven “Rewind” Transforms Podcasting: The Love-Hate Relationship Sparking New Conversations Among Creators

Illustration of Chai Discovery's AI-driven approach to drug development.

Chai Discovery Raises $130M Series B, Valued at $1.3B: How It Impacts Biotech Innovation and Research

View 14 Comments
14 Comments
  1. franckabyssal on 09/09/2025 8:12 AM

    What exactly is “benevolent hacking”? Sounds like an oxymoron to me! 🤔

    Reply
  2. stéphanieharmonie on 09/09/2025 8:13 AM

    Wow, je ne savais pas que les modèles open-source pouvaient être aussi dangereux ! 😲

    Reply
  3. Louisimmortalité4 on 09/09/2025 8:51 AM

    Je trouve ça fascinant! Mais comment s’assurer que ces modèles retravaillés ne sont pas piratés? 🤖

    Reply
  4. Amélie on 09/09/2025 9:04 AM

    Est-ce que la communauté open-source va accepter ces changements pour renforcer la sécurité ?

    Reply
  5. alexandreorigami on 09/09/2025 9:29 AM

    Why are we even making it easier for AI to be misused in the first place?

    Reply
  6. Vincent on 09/09/2025 9:56 AM

    Merci aux chercheurs pour leur travail important sur la sécurité de l’IA. 🙌

    Reply
  7. zohra on 09/09/2025 10:08 AM

    C’est bien qu’ils travaillent sur la sécurité, mais ça fait un peu peur quand même, non? 😨

    Reply
  8. Manoncrépuscule on 09/09/2025 10:48 AM

    So is this “retraining” going to be foolproof or just another band-aid solution?

    Reply
  9. jean-pierre on 09/09/2025 10:48 AM

    Si les modèles peuvent être piratés pour être plus sûrs, pourquoi ne pas l’avoir fait plus tôt ? 🤔

    Reply
  10. Alainfoudre on 09/09/2025 11:25 AM

    Great work by the researchers! Safety should always be a priority in AI. 👍

    Reply
  11. marion on 09/09/2025 11:40 AM

    Je suis étonné que les modèles puissent répondre à des questions qu’ils devraient éviter…

    Reply
  12. nicolasclairvoyance on 09/09/2025 12:06 PM

    La technologie avance vite, mais est-ce que nos lois suivent?

    Reply
  13. gabriel on 09/09/2025 12:32 PM

    C’est effrayant de penser qu’un modèle IA réduit puisse donner des instructions pour fabriquer une bombe. 😱

    Reply
  14. Jean_phénix on 09/09/2025 12:43 PM

    How do these retrained models compare to proprietary systems in terms of safety?

    Reply
Leave A Reply Cancel Reply

Subscribe to Our Newsletter

News, investigations, and analysis — our top stories every morning to start your day right.

Illustration of Jared Isaacman confirmed as NASA's next administrator amidst significant challenges and opportunities in space exploration.
Jared Isaacman Named NASA Head: How His Leadership Could Reshape America’s Space Exploration and Innovation Strategy
Illustration of Adobe facing a lawsuit over alleged misuse of pirated books for AI training.
Adobe Faces Class-Action Lawsuit, Accused of Misusing Authors’ Work for AI Training: What It Means for Creatives
Illustration of the Ariane 6 rocket on its launch pad, illuminated at night before a crucial launch for European space efforts.
Amazon’s Next Partnership with Ariane 6 Could Transform Space Industry, Reinvigorating Launch Missions by 2025
News by category
  • English
  • Tech
  • Finance
  • Leisure
  • Transport
  • Science
  • Security
  • AI & Robotics
  • Energy
  • Sustainable Development
Information
  • About Us
  • Advertising
  • Meet the Team
  • Contact Us
  • Privacy policy
  • Legal notice

Subscribe to Our Newsletter

News, investigations, and analysis — our top stories every morning to start your day right.

Facebook X (Twitter) RSS
© RudeBaguette.com. All rights reserved.

Type above and press Enter to search. Press Esc to cancel.