RIP, vector database

(turbopuffer.com)

205 points | by razin 4 hours ago

17 comments

  • gopalv 3 hours ago
    > This write amplification is large enough that our efforts to tune indexing throughput have started to hit diminishing returns.

    > don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.

    This is a direct parallel to how Postgres and Mysql built indexes.

    Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table.

    Postgres always points an index to a row-id within postgres which is an arbitrary value which changes on each update.

    Mysql, always assuming the storage engine is pluggable, points to the primary index entry and adds an extra indirection to the lookup.

    This means that you point the mysql index to a stable id, so unless you go update the primary key for a row, you won't have to update the indexes for all the attribute lookups you might have made to data.

    I don't do databases any more that much, but the design for NIMBLE file format has a lot of quirks which are relevant to this specific idea (wide tables).

    But the old Uber post about switching from Postgres to Mysql to prevent index amplification[1] is a direct mirror to this post.

    [1] - https://www.uber.com/us/en/blog/postgres-to-mysql-migration/

    • malisper 1 hour ago
      > Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table

      You are right that MySQL does better when you have lots of indexes, but I don't think the tradeoff is that the overall Postgres architecture is better with good schema design.

      Having secondary indexes point the primary key enables things like undo logging, which obviates the need for vacuums - vacuums being the most painful part of Postgres. On top of that your primary key index will be mostly cached so the cost of the indirection is much smaller than it may first appear

      • tomnipotent 51 minutes ago
        I think OP is just alluding to the fact that Postgres needs to do less work to go from secondary index to table data, since the tid is a direct pointer to the exact page and slotted entry while MySQL needs a b-tree walk.

        > primary key index will be mostly cached so the cost of the indirection is much smaller than it may first appear

        Not sure I follow. If it's in-memory you save having to read from disk, but you still have to walk the b-tree to go from PK to data.

    • phoghed 3 hours ago
      > mysql was optimized for a bad design

      TIL I should have been using mysql the whole time

      • woadwarrior01 3 hours ago
        Richard Gabriel's "Worse is better" vibes.
    • FLeXMurphy 2 hours ago
      I find it amusing people started quoting LLM output and are responding to it. Hopefully the original authors end up having the LLM respond back.
      • 0c3ca83 2 hours ago
        Many of the commenters on this site are also obviously LLMs. I'd imagine that quite a few of the entities quoting aren't necessarily people. Keep an eye on where they slide mentions of other products that a marketing team would like to promote.
        • FLeXMurphy 1 hour ago
          Personally I haven't seen it too often; the other aspect is that HN is a forum for startups to pitch shit to each other, so this has been happening with or without marketers (f.e. "i'm working on a similar thing").

          That LLMs are taking over the comments section is something that was already flagged, and Lobste.rs and others have started solving it by having gated registrations. HN should do this but it is unlikely to until it is too late.

          • alfiedotwtf 14 minutes ago
            I don’t get it though… what’s the point of using an LLM just to comment here? Like what does it gain the person doing it
          • awesome_dude 1 hour ago
            Sorry, how does a gated registration stop someone creating an account, then handing it over to an LLM?

            Serious question, am I misreading what's being said?

            • lukan 1 hour ago
              It stops the automation part. One LLM spammer can be dealt with.
              • unglaublich 33 minutes ago
                One can focus on generating 10 other accounts before it's flagged. Now you have 10 others to deal with.
                • lukan 21 minutes ago
                  But that does not work on lobsters (or anywhere else with gated registration), because you cannot just create 10 accounts. You can create 1 account after 1 invite. If you abuse, you loose the invite (and maybe the person inviting you the right to invite others).
          • senderista 1 hour ago
            lobste.rs has always been invite-only.
    • SigmundA 55 minutes ago
      MSSQL (Clustered Indexes) and Oracle (Index Organized Tables) among others let you chose because there are advantages and disadvantages for different situations.

      Not having true clustered indexes in PG is something I miss coming from MSSQL, it helps performance when the majority of access is always primary index avoid indirection from index lookup then tuple lookup and it also saves space if its the only index.

  • real_faxenoff 4 minutes ago
    I’m developing a local “code graph mcp tool” (not yet published) and followed a similar path, though I may have been able to go further since I have fewer vectors in my database (even on projects with 50M LOC).

    At first, I tried all those popular vector databases and was disappointed with their performance. In the end, the best and fastest solution turned out to be building a multi-database system on SQLite, compiled with everything related to multi-client operations removed. Only exclusive mode was left. Everything is as binary as possible. The index is completely separate — an IVF with pre-training — and is built on the GPU (250K vectors are built, processed, and saved in 4 seconds). Right now, my biggest problem is frequent data changes, and I need to implement optimizations to reduce recalculations.

    So far, I haven’t seen any vector database implementations that are heading in the right direction. Maybe only Lancedb looks promising, but it’s too heavy for my needs.

  • gk1 3 hours ago
    Vector databases were always more about retrieval than either vectors or data storage. But the term stuck all too well and companies held on to it a tad too long. Sorry :)
    • tveita 1 hour ago
      That's just a search engine, but then you're competing with traditional players like Elasticsearch and Vespa who all have built-in vector support by now, and you have to compete on attributes like price, performance, features, and who can mention 'AI' the most times on their web page.
  • tschellenbach 1 hour ago
    AI has some of the craziest up and down cycles of tech I've ever seen
  • Tsarp 2 hours ago
    I've really liked lancedb for similar use cases. Not just that it is OSS. But Lance treats ANN as a secondary index similar to what turbopuffer v3 does. Rows sit in fragments, and the vector index never moves them.
  • marekgalovic 2 hours ago
    > The problem with a vector primary index

    We've realized this a long time ago at TopK and built a flexible serverless search engine from scratch. Supports dense/sparse vectors, late interaction, lexical search, indexed regex, filtering, and custom scoring in one query.

    - https://www.topk.io/blog/vector-dbs-are-the-wrong-abstractio... - https://www.topk.io/blog/topk-embed-v1

  • ironqcold 1 hour ago
    I'd want to see p99 at 1k+ QPS on the same scale
  • drewlanenga 3 hours ago
    the multi-vector duplication thing makes sense, copying every attribute once per vector explodes quickly. what's the new primary index?
    • benesch 1 hour ago
      An automatically generated internal ID: (segment ID, doc ID). The user-provided primary key (the field called `id` in the document) turns into a secondary index at the storage layer.
  • orliesaurus 2 hours ago
    Waitint for the CEO of Qdrant to step in
  • ActorNightly 2 hours ago
    Im not full read up on RAG pipelines, but has anyone ever tried to make the database a neural net itself? I.e get rid of any sort of traditional databases, and then you basically just have some sort of autoencoder?
    • tyromaniac 42 minutes ago
      In some sense there's probably a database compression scheme that does something similar. Usually people care too much about fidelity
  • sreekanth850 3 hours ago
    I find very little reason to use a pure vector database for enterprise retrieval. We built an enterprise retrieval engine on top of a SQL database with native vector support, and the flexibility is something we cannot ignore. Vector similarity is just one query primitive alongside full text search, filters, joins, ordering and normal relational predicates. Tenant/app/collection isolation becomes part of the query itself. ACLs, document versions, categories, metadata constraints and temporal filters are ordinary predicates rather than something you have to bolt onto a vector store. SQL is already going to be part of almost any enterprise system. Adding a separate vector database introduces another moving part and syncing two system whenever you update your data is the most difficult thing to get right.
    • ijidak 1 hour ago
      Which database did you use?
      • sreekanth850 49 minutes ago
        We use cratedb. but now, clickhouse, starrocks all hve vector.
    • polynomial 1 hour ago
      [dead]
  • blakeashleyjr 3 hours ago
    This sounds like the Postgres vs. InnoDB argument 10 years later. Postings pointed at physical location (the ANN slot), so every SPFresh rebalance rewrote every index touching that doc. InnoDB solved this by pointing secondary indexes at the PK and eating an extra lookup on read. Curious what that extra lookup costs you when it's an S3 GET instead of a B-tree hop.

    "Updating one vector can move hundreds of attributes and their indexes" is basically Uber's 2016 Postgres write amplification post, but for search. Same fix too: stop pointing indexes at where the row lives.

    So ANN becomes a secondary index that points at a doc ID, and vector search now needs a hop to complete. Do clusters keep their own copy of the vectors so the search itself stays local, and only result fetch pays the indirection? Otherwise cold p99 seems like it gets worse.

    • alfiedotwtf 11 minutes ago
      I haven’t read it, but “ stop pointing indexes at where the row lives” sounded interesting. So if not the row, what does the index point to instead?
  • OutOfHere 3 hours ago
    It would be nice to have a page that actually loads. This one doesn't. RIP.

    UPDATE: It loads now, but it didn't when it was first posted. Traffic load on the server does matter.

    • throwawy0352 3 hours ago
      Loads really fast for me. (MacBook Air, average internet)

      If you still have issues, try https://web.archive.org/web/20261001100105/https://turbopuff...

      • wilj 3 hours ago
        It has a pagespeed insights score of 55 and noticeably sluggish on my m3 max.

        And what's with the throwaway account for this one comment? Is this becoming reddit with throwaway shills now?

        • phoghed 3 hours ago
          fucking shills, making helpful comments and promoting seemingly nothing, what's this place coming to?
        • throwawy0352 3 hours ago
          Yes, I get big money from the Internet Archive to promote their services. It's the new scheme that shills like me go for.

          The reason is that I have no account on HN and rarely comment. I create a new account a few times a year because I don't remember or care about my previous account.

          I could have made an account named john2026 and you would not think twice. Instead, I let people know upfront what type of account this is. Quite the opposite of what a true shill would do.

          I got a Lighthouse score of 99 in Chrome. Believe it or not, I won't spend more of our time on this. (relevant XKCD: https://xkcd.com/386/ )

          First Contentful Paint 0.7 s

          Largest Contentful Paint 0.9 s

          Speed Index 0.7 s

          It makes a lot of requests, and some are stopped by my ad blocker, but most of them don't seem to make an difference. It is almost instant from my point of view. I disabled the ad blocker and didn't notice any visual difference.

    • syndacks 3 hours ago
      loads just fine on my $10k laptop with 10g internet here in NYC
      • alexjplant 3 hours ago
        Takes 11 seconds to load on Firefox on Linux with 3G-level throttling enabled in Dev Tools.
      • uproarchat 3 hours ago
        Also loads fine on my beater in the sticks :)
      • OutOfHere 2 hours ago
        Do you actually think that 10G makes pages load faster than 1G or even 100M? It doesn't. The blocker was most likely on the source server, not on your side.
        • WarcrimeActual 2 hours ago
          You're the reason the /s tag has to exist.
          • phoghed 1 hour ago
            The onion is more valuable when there are people to eat it, we should be thanking him
    • swedishPerson1 3 hours ago
      [dead]
    • jasonmp85 3 hours ago
      [dead]
  • vhiremath4 1 hour ago
    AI Slop. Will not read.
    • gravitronic 1 hour ago
      The article?

      turbopuffer is founded by some of the smartest people I ever worked with in past jobs. I strongly doubt they used an LLM in the writing of this article.

      • _peregrine_ 1 hour ago
        confirmed - we still write by hand
    • dolebirchwood 1 hour ago
      Are humans who use em dashes that intimidating to you?
  • childintime 2 hours ago
    Is it time to kill the database and replace it with a LLM optimized compiled version that simply implements the required API directly in (Rust) code, without any dynamic overhead? It probably will still be based of off a base design or a base file format.

    Ultimately this system will encompass the whole OS, of course, but the DB might be the best place to start.

    • tyre 1 hour ago
      You mean get rid of Postgres and build bespoke database-esque systems for every use case?

      If so, then no. It is not time for that.

      • cogman10 1 hour ago
        Yeah, I'm struggling to come up with a really good time for that.

        The best I got is if you are trying to do an old-school style video game asset/save game storage. But even then, the value in just using sqlite or even parquet is really high.

        There's so many really good data formats that deciding on a new one at this point seems pretty silly. Particularly because what you sign up for when you make a new one is losing any and all tools that could be used to work with and diagnose that data.

    • nemothekid 1 hour ago
      Instead of a database, the LLM will expose an api endpoint and build a database on demand?

      That's interesting. Maybe to decrease latency the LLM could "cache" it's build of it's database and reuse in between instances. It could host this artifact on a "hub" of git trees and then any new use cases that come up, can be added to this git tree. Then it can possibly be reused in different use cases.

    • bijowo1676 1 hour ago
      SQLite already exists and some people use it
    • dymk 1 hour ago
      Is it time to get rid of hammers and replace them with swiss army knives?
      • pessimizer 7 minutes ago
        We could build a new hammer for each nail!
    • eatonphil 1 hour ago
      I have seen this happening already at two different companies. And I'm also doing it as well. Particularly for search indexes where there's no risk of data loss.
    • Ostatnigrosh 1 hour ago
      Unless you're tigerbeetle and want to handroll every single thing you do lol