26 comments

  • xrd 1 hour ago
    About twenty years ago, I was taking a flight back from Rio de Janeiro, Brazil to the US. In the middle of the night the pilot got on the loudspeaker and said "hi! Having some engine trouble, so we are landing in Manaus."

    Manaus is in the middle of the Amazon.

    Needless to say, a bit scary to hear that, but we landed without issue.

    They told us we had two choices: the nice hotel with a shared room, or the lesser nice hotel with no roommate. I chose the latter. When we go there, they said, "oops, sorry, short on rooms!" So I had a roommate.

    Wandered around Manaus, took a skiff out on the Rio Negro. Saw pink river dolphins. A little boat approached us and a kid handed me a sloth, and then demanded I return it with a twenty dollar bill.

    The airline got us another plane 24 hours later. Made it back to the US safely.

    A few weeks later, the airline reached out and said "Here is $100 for your trouble."

    I declined to take that offer. I had missed several business meetings that cost me actual money. I couldn't donate blood for years because I had been to the Amazon and was tagged a malaria risk.

    During the many arguments with the airline I threatened to take them to small claims court.

    I got a really strange response over email which I clearly wasn't supposed to see. A representative from that airline was asking internally if they could put me on the no-fly list. That was really chilling.

    But, this is the kind of information I'm worried about when a vendor sells my data. If Google wanted to sell a product to the airlines that offered to keep annoying people like me from purchasing flights, they could do that with that email chain. I'm skeptical it'll be wiped correctly. Isn't my poor writing style basically my signature? How do you wipe that?

    • elric 58 minutes ago
      I fully share your concerns. And I don't understand how apparently tons of Teams and email conversations can be archived and sold without any kind scrutiny. How can such data be sold without the consent of all involved parties? What gives Google the right to use it to train LLMs? Is that just a way of washing away the legal protections?
      • Ekaros 38 minutes ago
        Makes one appreciate living in place with sufficient constitutional protections against this sort of stuff. Even for work stuff selling this info wouldn't fly in some parts of the world.
        • woadwarrior01 22 minutes ago
          If you're referring to GDPR, companies routinely evade such protections using "informed consent" / "legitimate interests" loopholes. The big ones get caught once in a while, get a slap on the wrist and continue to do whatever they were doing before, albeit with more safeguards.
      • warkdarrior 36 minutes ago
        The party owning this data (Spirit Airlines) is consenting to the sale. Employees and customers of Spirit consented when they started employment and did business with Spirit, respectively.
        • Leynos 4 minutes ago
          This is why the GDPR (and to a lesser extent the CCPA) is a good thing. The data was supplied for a specific purpose. The handler of the data should have to obtain further consent if they wish to use it for another purpose.
        • alberto-m 18 minutes ago
          Did they consent? Just because one receives a letter it doesn't mean they “own” it, much less that they are entitled to publish it at their leisure. If Spirit were active in any country with GDPR-style laws, the seller of these data would be most likely investigated.
          • hdgvhicv 13 minutes ago
            America believes in freedom for large companies to take personal data and make it their own, rather than individual feeedom
        • MagicMoonlight 29 minutes ago
          [dead]
    • PaywallBuster 1 hour ago
      your personal site SSL cert expired 10 days ago btw
      • dwedge 53 minutes ago
        I love the irony of you checking them out for more information in response to a comment of them being worried about who reads their data. Nothing wrong with it, just make me chuckle
      • xrd 36 minutes ago
        Doh, thanks!
    • snickerbockers 39 minutes ago
      So what happened to the sloth??? Don't bury the lead, man!
      • xrd 31 minutes ago
        The sloth was returned to his owner and I did tip him. That kid is probably still prowling the Amazon (as an adult now), looking for sucker tourists like me.
    • testing22321 24 minutes ago
      Did you end up getting more than $100?
    • warkdarrior 42 minutes ago
      Manny retail industries already share lists of "troublesome" customers (trouble = anything from too many returns to lawsuit-happy to friendly fraud). Not sure this is a new concern..
    • jefftk 50 minutes ago
      I think you might have missed the deidentification piece?
      • xrd 43 minutes ago
        Not trying to be snarky, and perhaps it wasn't well stated, but the last paragraph I said I'm concerned about identification via my writing style. If they have my emails, they would have my writing style. It doesn't have to be tied to PII there, they can cross reference it with my blog. I'm speculating because I read that you can identify people by a few sentences of their writing.

        "Deidentification" seems really murky and imprecise at best.

      • fileeditview 44 minutes ago
        You have to trust that this really "deidentifies". Time and time again it was shown, that the measures taken were not enough to anonymize.

        E.g. the parent wrote that he fears, he could be identified by his writing style, which is totally plausible. How would you "deidentify" this?

        • piva00 35 minutes ago
          Even if they follow to the letter a deidentification process, Google and Meta have so much data about individuals that re-identification shouldn't be very hard for the majority of airline passengers' data they put their hands on.

          Of course, takes a lot more effort than not doing proper deindetification in the first place but if they wanted to appear like caring about data privacy they still have enough data points to correlate the sets later on (and/or over time).

      • reaperducer 45 minutes ago
        No such animal.
      • Leonard_of_Q 40 minutes ago
        I have a bridge for sale, hardly seen use, pay me ${money} and you can collect it in New York City. Interested?
  • js2 1 hour ago
    > Google bought itself 100 million emails and 500 million items from Microsoft Teams, 17 million OneDrive files and 20.5 million items from SharePoint. The search giant also now owns over 30 million recorded customer service calls, and more than 15 million customer service chat records. 600,000 ServiceNow tickets are another element of the collection, along with 13.7 million active emails addresses from Oracle’s Responsys marketing application, and details of 11 million sales of in-flight Wi-Fi services.

    > There’s also operational data in the trove, describing over 763,000 flights, five million crew pairings, more than 1.2 million fuel slips, and records describing purchases of 787,452 parts.

    > Google has reportedly said it bought the data to improve its AI services.

    Gives "this call is being recorded for training purposes" new meaning.

    • dgellow 34 minutes ago
      Is there anything that can legally be done against this? It feels like a breach of consent. Like, it cannot be that when one accept their voice to be recorded for _human_ training they also accept it to be recorded for LLM training
      • hdgvhicv 6 minutes ago
        You’re years too late.
    • echelon 59 minutes ago
      "This call is being recorded so that Gemini can decide which purge wave to assign you to. Obedient humans will be carried over for further cycles until no longer needed. If you are scheduled for termination this cycle a disposal representative will be with you shortly."

      I kid, but...

      It's probably the precursor to insurance denials and job screening.

      I got banned from r/technology a few weeks back for decrying tracking in AI content. The community was piling on saying it was okay because it removed AI content or made it easy to spot. I made the counter argument that watermarks would find their ways into everything and eventually be bound to attestation. The mods didn't like that. (Yet another structural problem with the lack of p2p self-service town squares.)

      The socials are training the next generations for broad acceptance.

      • TeMPOraL 45 minutes ago
        > I made the counter argument that watermarks would find their ways into everything and eventually be bound to attestation.

        Yup.

        Elsewhere in another front page thread today: "oh but apps blocking screenshots because of 'sensitive content' can be bypassed by taking a photo of your screen with another phone".

        Any tech-savvy person with two brain cells reading this and that: "gee, I wonder if the same magic imperceptible watermark that survives multiple rounds of cropping and printing and scanning, that's used to tag AI-generated content, could also be used to tag sensitive data, or ads, or which app is rendering it on screen, and then the camera app could refuse photographing it...".

        I don't know why people don't see that AI watermarks are DRM, and DRM is universal, and there are many clients...

      • dboreham 36 minutes ago
        Finally we know how they assign people to either that A Ark or the B Ark.
  • ronbenton 1 hour ago
    > 600,000 ServiceNow tickets are another element of the collection, along with 13.7 million active emails addresses from Oracle’s Responsys marketing application, and details of 11 million sales of in-flight Wi-Fi services.

    I really doubt all this stuff was “de-identified”

    • imglorp 51 minutes ago
      I don't see how it's possible any more, when correlated against all the various other data sources. And a record that might be unidentifiable now might become unique with more correlated data sources.
    • chii 52 minutes ago
      > de-identified

      De-identified but far from useless.

      as an example, they can remove the names off these sales data, so you can't identify who purchased what items. However, the purchaser would be identified by some sort of number, and you would be able to extract information about purchasing habits, and aggregate these habits into usable information for advertising purposes (like targeting and profiling).

      And that's before AI training for LLM purposes.

  • pm215 55 minutes ago
    I see from the court PDF that the process here involves Spirit giving the data to a "Deidentification Agent" (a third party firm that Google selects and pays for) who is responsible for stripping out things that would link data to any particular person before passing the data on to Google. Is that a standard thing, such that everybody in this transaction would have said "yes, put in the usual clauses about deidentifying the data" and multiple firms offer this service, or is it something that they custom-specified for this "we want the data for AI" transaction?

    (The PDF mentions "the standard for deidentification set forth under the California Consumer Privacy Act", which suggests this is all pretty well legislatively understood.)

    • peyton 44 minutes ago
      Seems the answer is “no” to the first part of your question. From the filing:

      > For example, one initial bid requested certain customer list information; however, by the first round of the Auction, the most competitive bidders had agreed to bid on an asset schedule that expressly excluded PII.

  • hn_submit 7 minutes ago
    I wonder if the owners of YCombinator have sold all comments to Big Tech for A.I. training.

    And how long before Google and Microsoft add to their T&Cs that all your anonymized email will be used to train their A.I.?

    • bitmasher9 2 minutes ago
      Are all of the messages public and crawlable?
    • Forgeties79 3 minutes ago
      I don’t know why any company would pay. It can’t be that difficult to scrape this site and they already did it indiscriminately for years, violating laws and taking down public libraries and other public resources with no regard for their impact.
  • andy99 16 minutes ago
    I wonder how they will use the data. If it was me I’d try to build a simulation of an airline, and then use it as an agent training environment. It really depends on the exact nature of the data what kinds of agents you could train, but maybe customer support (imo the worst AI use case) that are more empowered to make changes, or something for making more autonomous calls when recovering from irrops? Could be some cool’s stuff if a little niche, I hope they share / publish something and it doesn’t just disappear into a void.
  • Ekaros 1 hour ago
    Anyone else somewhat weirded by current state of affairs that this sort of information is valuable enough to even bother selling... And that it actually happens... It feels like some societies are in really weird place.
    • budman1 3 minutes ago
      How does this have value? Is any and every sentence in an e-mail considered 'fact' and thus to be fed into the AI?

      90% of e-mails and Teams communications are inane. Polite banter, "thanks for taking care of that, I appreciate it" "please route the forms to Janet this week because Bill is on vacation" "unit will be un available until the parts come in" . I can't see the intrinsic fact value of this kind of communication without screening it. And after screening, the gold nuggets would be minimal.

    • greggoB 1 hour ago
      I am very much weirder out by it, yeah. Seems some societies are just excessively desperate for some kind, any kind, of fuel for economic growth, to the point this is where attention is now. The term "post capitalism" being thrown around feels less ridiculous than it did in years gone past.
      • usrusr 28 minutes ago
        Any kind of fuel for giving active investors that FOMO tingle which then forces the steamroll of index funds to blindly follow.

        I guess the appropriation "any sufficiently advanced stock market is indistinguishable from entertainment" doesn't quite stop at equating the trade floor with a casino. At some point, entertainment also becomes the modus operandi of corporations.

  • blitzar 1 hour ago
    The headline is a little on the nose. Nice try but it isnt going to hit the levels of "Headless body in topless bar".
    • isoprophlex 1 hour ago
      a vegetarian dinosaur, called "the quick bandit", eats shoots and leaves! no idea where he got his name.
  • wewewedxfgdf 6 minutes ago
    I can make them millions of emails and I'll charge them only $2 million not $10 million.
  • sethammons 1 hour ago
    So the AI service agent can be just as bad as Spirit's service was.
    • genxy 24 minutes ago
      Google can now stamp out copies of autonomous corporate minds that are clones of Spirit. Haunting.
  • KORraN 1 hour ago
  • balderdash 1 minute ago
    Great - now spirit will be the “model”, could you pick a worse example?
  • zf00002 17 minutes ago
    The idea that they got an archive of my coworkers tickets that just say "its broke", is amusing.
  • genxy 26 minutes ago
    So they didn't even have to build the torment nexus, they just bought it. Only bad can happen.
  • Forgeties79 6 minutes ago
    That title is a bit much. I get they’re going for “wordplay” but at first glance I thought they were buying the data from crashes flights…?
  • lovetocode 33 minutes ago
    I can’t help ask but what? They say it’s for training their models. On what? One of the most horribly run airlines to ever exist?
    • TeMPOraL 20 minutes ago
      > On what? One of the most horribly run airlines to ever exist?

      On real life data on operations of a real large company.

      Internally, most big companies are probably just as big of a mess, if not worse. But you can't get that data easily.

    • Mistletoe 16 minutes ago
      “Gemini, do the opposite of everything in the Spirit archives.”
  • steveBK123 1 hour ago
    Maybe they are building a social credit score system
  • dec0dedab0de 58 minutes ago
    funny, spirit was the only big airline without a crash
  • cmiles8 1 hour ago
    “This call is being recorded for quality assurance, and to give us more assets to sell in bankruptcy.”
  • everyone 1 hour ago
    How the fuck is that even remotely legal? ... "deidentified" my ass.
    • embedding-shape 1 hour ago
      Ah, but Google promised to remove PII they found in this deidentified dataset, so worry not.

      > If you’ve flown Spirit and worry that Google will soon know about a testy conversation you had with the airline’s call center, you’re being told not to worry. The court filing says the data was deidentified before being put on sale and Google has promised to scrub any PII it finds in the trove.

      • akoboldfrying 1 hour ago
        I basically agree, but I'd also say: Every Gmail user has already accepted such a promise as sufficient.
        • phatfish 50 minutes ago
          I wouldn't be surprised if Gmail data has far more access restrictions internally at Google than this auction bought dump.
      • sscaryterry 1 hour ago
        Ah, trust me bro :)
    • noir_lord 1 hour ago
      You know when we (they) tell you not to do any personal computing on work devices/systems and to keep your devices completely separate from work ones.

      Yeah this (and lawsuits/investigations) are why, the employer owns the data, in some contexts (like this one) it can become an asset (or a liability) but in either case it's not yours.

      Of course that only gets you part of the way there anyway see Twitch recently opting in all users by default to mined for AI and only adding an opt out after backlash with a quote that was so on the nose it made me stop "If we'd have asked them to opt in, they wouldn't have opted in" (paraphrasing but it was that blunt).

  • Razengan 51 minutes ago
    I would love it if there were services where I could let them see everything I do, including when I poo and wank, if they just directly paid me for it.

    No I don't want to just use your enshittified service for free. Fucking pay me and watch me all you want :)

    • xyzelement 10 minutes ago
      We don't want your data if you are getting off on it.
  • evek 50 minutes ago
    Tangental, but can’t wait for automated blackmail from crawlers continuously digging through my digital footprint. /s

    Martha Wells hit it nicely in The Murderbot Diaries.

  • Taikhoom10 1 hour ago
    [flagged]
  • radres 1 hour ago
    they wrote a shit article while trying to make some airline puns
    • stuaxo 1 hour ago
      Woke up on the wrong side of the bed?
  • 398642258909 1 hour ago
    Each Register headline is worse than the last one.