• w3dd1e@lemmy.zip
    link
    fedilink
    English
    arrow-up
    16
    ·
    7 months ago

    Firefox kept crashing on me a few days ago. Decided to run MemTest86 and sure enough. Bad RAM.

  • flamingo_pinyata@sopuli.xyz
    link
    fedilink
    English
    arrow-up
    8
    ·
    7 months ago

    This is how dev humblebrag sounds like.
    Our app is so stable only random hardware events like bitflips can crash it.

    • grue@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      6 months ago

      LOL, nah, Firefox isn’t that stable. If 10% of crashes were caused by bad RAM, it means 90% were still caused by something else.

      (My install regularly gets a memory leak that eventually makes my system unusable, BTW. I don’t think it’s necessarily the fault of Firefox itself – more likely Javascript running in tabs, maybe interacting with an extension or something, and some of the blame goes to the kernel’s poor handling of low memory conditions – but it’s definitely not “dev humblebrag stable” for me.)

  • Toes♀@ani.social
    link
    fedilink
    English
    arrow-up
    7
    ·
    6 months ago

    I used to be a part of an anticheat dev team and we discovered that this was a common problem back in the Windows XP era.

    We added a routine to check the memory addresses used after a crash and notified the user if we suspected hardware failure.

    At the time we suspected unstable overclocks because the metrics showed us the computers affected were typically overclocked as well.

    • llii@discuss.tchncs.de
      link
      fedilink
      English
      arrow-up
      4
      ·
      7 months ago

      When I upgrade my home server I would like a low-power system with ECC RAM. I hope it will be financially viable in the future.

      • tal@lemmy.today
        link
        fedilink
        English
        arrow-up
        3
        ·
        edit-2
        7 months ago

        The problem is that ECC is one of the things used to permit price discrimination between server (less price sensitive) and PC (more price sensitive) users. Like, there’s a significant price difference, more than cost-of-manufacture would warrant. There are only a few companies that make motherboard chipsets, like Intel, and they have enough price control over the industry that they can do that. You’re going to be paying a fair bit more to get into the “server” ecosystem, as a result of that.

        Also…I’m not sure that ECC is the right fix. I kind of wonder whether the fact is actually that the memory is broken, or that people are manually overclocking and running memory that would be stable at a lower rate at too high of a rate, which will cause that. Or whether BIOSes, which can automatically detect a viable rate by testing memory, are simply being too aggressive in choosing high memory bandwidth rates.

        EDIT: If it is actually broken memory and only a region of memory is affected, both Linux and Windows have the ability to map around detected bad regions in memory, if you have the bootloader tell the kernel about them and enough of your memory is working to actually get your kernel up and running during initial boot. So it is viable to run systems that actually do have broken memory, if one can localize the problem.

        https://www.gnu.org/software/grub/manual/grub/html_node/badram.html

        Something like MemTest86 is a more-effective way to do this, because it can touch all the memory. However, you can even do runtime detection of this with Linux up and running using something like memtester, so hypothetically someone could write a software package to detect this, update GRUB to be aware of the bad memory location, and after a reboot, just work correctly (well, with a small amount less memory available to the system…)

        • AA5B@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          6 months ago

          I wonder if ai can actually help here. As the industry abandons consumer hardware in favor of datacenter equipment to profit from the ai bubble, perhaps ecc memory will become cheaper

        • grue@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          6 months ago

          There’s no real good reason that all RAM shouldn’t have been ECC since decades ago. It doesn’t actually cost much more to implement. The only reason it isn’t, as tal’s reply mentioned, is artificial price discrimination.

    • thebestaquaman@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      7 months ago

      You can’t effect the number of bit flips your users hardware has, but you can affect how often buggy code corrupts their memory or otherwise crashes your program.

      Let’s say any app will crash about once a year on my machine due to a bit flip. If the app is crap and crashes hundreds of times for other reasons, the bit flip is irrelevant. If the app is robust enough that the bit flip accounts for 10 % of the crashes, that basically means the app is pretty much never crashing due to poor code.

      • MoogleMaestro@lemmy.zip
        link
        fedilink
        English
        arrow-up
        2
        ·
        7 months ago

        That’s the way people should be looking at it. It basically means hard crashes are extremely rare in the firefox ecosystem.

        To be fair, I can’t remember the last time a browser crashed on me in general.

        • caschb@lemmy.world
          link
          fedilink
          English
          arrow-up
          1
          ·
          6 months ago

          I’ve had Safari of all things crash on me a couple of times. Still, not enough to actually be disruptive.

    • Kairus@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      7 months ago

      You’re assuming that app quality is constant. But if I made an app that crashes on launch, I can confidently say 0% of those crashes would be from bitflips.

      Firefox isn’t special in some way that could cause bitflips, but it’s 1) where this data was collected (and why this post isnt talking about some other product) and 2) speaks to the quality of FF, because crashes are rare enough for bit flips to be a significant crash factor.

      The takeaway is that for the FF team, and anyone using ram (everyone), bitflips are more common than expected

    • r00ty@kbin.life
      link
      fedilink
      arrow-up
      1
      ·
      7 months ago

      No, they’re saying Firefox uses so much ram they’re far far more likely to be a victim!

      • grue@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        6 months ago

        Laughs in Memory: 46.84 GiB / 62.72 GiB (75%) with (probably) several hundred tabs open

    • Deestan@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      7 months ago

      As a long time Firefox user, I believe Firefox sees orders of magnitude more RAM issues than other apps because it is using orders of magnitude more RAM than other apps.

    • JensSpahnpasta@feddit.orgOP
      link
      fedilink
      English
      arrow-up
      0
      ·
      7 months ago

      It would be interesting to see how this works in Chrome. I would guess that it could be the same - people tend to leave their browsers open with hundreds of tabs and will never reboot their laptops. If you play a random game for 2 hours, bit flips shouldn’t be a problem. But if you keep your browser open for weeks or months with hundreds of tabs, that may cause problems.

      • Jarix@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        arrow-down
        1
        ·
        6 months ago

        … I can’t imagine having a browser with hundreds of open tabs. That would tend me of the old days of Netscape Navigator and all the popups and browser add on cancer.

        Ahh the nostalgic days of the early Dotcom era. I sometimes miss you geocities

    • tal@lemmy.today
      link
      fedilink
      English
      arrow-up
      0
      ·
      7 months ago

      Anecdotal evidence, but I had both a 13th gen and 14th gen Intel CPU with the bug that caused them to over time, destroy themselves internally.

      The most-user-visible way this initially came up, before the CPUs had degraded too far, was Firefox starting to crash, to the point that I initially used Firefox hitting some websites as my test case when I started the (painful) task of trying to diagnose the problem. I suspect that it’s because Firefox touches a lot of memory, and is (normally) fairly stable — a lot of people might not be too surprised if some random game crashes.

      • SleeplessCityLights@programming.dev
        link
        fedilink
        English
        arrow-up
        1
        ·
        6 months ago

        I had to turn down my block multiplier so that I could play Unreal Engine games. I would suspect that would extend the lifetime. After the underclock I have perfect stability.

    • Jarix@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      arrow-down
      1
      ·
      6 months ago

      How so?

      Didn’t it just highlight how stable the software is?

      I assume bitflipping crashes most softwares. If your software is so stable that hardware errors that effect everyone equally(which may be my erroneous assumption I’ll admit) then it is staying that if Firefox is crashing on you, it might be time to run some diagnosis on your hardware.

      A litmus test as a browser

      • xxce2AAb@feddit.dk
        link
        fedilink
        English
        arrow-up
        2
        ·
        edit-2
        6 months ago

        Fair question. I find it unnerving, because there’s very little a software developer can meaningfully do if they cannot rely on the integrity of the hardware upon which their software is running, at least not without significant costs, and ultimately if the problem is bad enough even those would fail. This finding seems to indicate that a lot of hardware is much, much less reliable than I would have thought. I’ve written software for almost thirty years and across numerous platforms at this point, and the thought that I cannot assume a value stored in RAM to reliably retain it’s value fills me with the kind of dread I wouldn’t be able to explain to someone uninitiated without a major digression. Almost everything you do on any computing device - whether a server or a smart phone relies on the assumption of that kind of trust. And this seems to show that assumption is not merely flawed, but badly flawed.

        Suppose you were a car mechanic confronted with a survey that 10 percent of cars were leaking breaking fluid - or fuel. That might illustrate how this makes me feel.

        • Jarix@lemmy.world
          link
          fedilink
          English
          arrow-up
          3
          arrow-down
          1
          ·
          6 months ago

          Hmm thanks, also please massively digress if you would like to.

          I interpreted it like 10% is a lot if it’s 10% of a million. That 100,000. So if there’s a million things that crash Firefox that’s a high number.

          If Firefox only crashes 10 times a year because it runs that well, 10% or that 1 time it crashes from a bitflip is impressive that the rare bitflip takes up such a high percentage of total crashes because Firefox just doesn’t crash very often.

          If your dread is found to be justified that won’t be too surprising, to me, if hardware is getting made less reliable these days thing. Enshitification being the norm, and tech being in everything nowadays

          We obviously need more context from Mozilla, but this could be a canary in the mine type situation.

          But it would be kind of neat if Firefox became something of a reliable test for bitflipping unintentionally

          • xxce2AAb@feddit.dk
            link
            fedilink
            English
            arrow-up
            2
            ·
            6 months ago

            I agree, and there are a number of other biases to consider. Here’s some I can think of:

            • Firefox will mainly be running of desktops, laptops and smartphones. I would expect QA to be significantly better for this type of device than, say, consumer grade routers or TV boxes. But more concerning to me is stuff like cheap ATMs, industrial control systems (although Siemens have great QA) and elevator control systems etc. Infrastructure, not consumer toy, and Mozilla obviously aren’t the right people to say anything about the state of any of that.
            • While Mozilla is currently estimating approximately 200 million installs, some of those - especially on Linux - will have disabled telemetry. I know I do. With that said, I can’t recall the last time I had a FF CTD (crash to desktop) but I suspect when I did, it wasn’t even a bug but an OOM (out-of-memory) kill because I was browsing on something like a 2Gb RAM micro-portable with insufficient swap. FF is one impressively stable piece of software these days.
            • Firefox usage is not evenly globally distributed, and I have no way to reliably assess whether FF has a larger or smaller proportional usage in regions that may rely more on older or refurbished hardware, which I would expect to have higher HW error rates (although I cannot prove that either - I can’t find any good public aggregate data for RAM MTBF trends over time, but I’d be very interested if somebody else knows where to find authoritative answers on that).

            (Un)fortunately, this may be the most Mozilla can provide in terms on insight. Their users tend to be particularly sensitive of perceived or practical privacy violations, so I understand - and appreciate - their caution in gathering data.