Video yükleniyor...
Video Yüklenemedi
My cofounder and I are making an FPGA-accelerated server for high-memory high-bandwidth workloads. We're looking for one or two companies to partner with; we'll do the work to port your application to our hardware. Please DM me if you've got a tricky workload!
79,906 görüntüleme • 11 ay önce •via X (Twitter)
36 Yorum

You can DM me, email me at [email protected], or schedule a call here: I'd love to chat about what you're working on, or if you know anybody else I should reach out to, please let me know!

Video encoding is a tricky workload. There are only a handful of FPGA encoders left.

Id like to interview you about the importance of FPGAs for our future, what do you say?

I'll DM you!

I’ve always believed FPGAs could be a strong fit for the future of AI computing, I even almost started a thesis on it. Why do you think they haven’t seen broader adoption?

Going to be at SC25?

I'm considering it! Will you be there?

Always. One of the hottest events of the year.

actually, i gave up right away on moving the code to the gpu, because the memory bandwidth of gpus is not that different from cpus *if you miss to main memory all the time*

Not a company but Leskovec is the Graph ML guy at Stanford and I think he regularly works with 10TB datasets and non-standard algos.

do you have some docs or a blogpost on the workload porting part of the operation? i'm super curious about it

Not yet, but we probably will at some point in the upcoming months! I think it would be fun to talk more about that.

for sure. thanks

Hey, that looks really interesting! (Tiny unsolicited advice: you might want to either unpin or release a playable version, because potential partners might worry you're the kind of founder who never ships 😅 not that I’d know anything about that!)

Why not include HBM through the relevant Xilinx or Altera parts?

Yes, very reasonable question. It's a complicated calculation whether HBM pencils out as better for any given workload. The design is extremely modular, and I think we are likely to add some HBM to some of the modules at some point, for workloads that benefit from it.

I love this idea.

@jaysidd fpga servers

What kind of processor does it use? Is the bandwidth before or after hardware compression? Is rand read as fast and what is the optimal block size?

The host Linux system is an AMD Epyc platform. The FPGAs are from Xilinx. These are raw speeds of the various interfaces, not compressed speeds. The flash speed is that we have ~160 separate NVMe interfaces, so re: rand reads, picture the performance characteristics of ~160 SSDs.

so somewhere in between positron atlas and titan. 2u 2-3kw aircooled, ~$100k per server? they have a simple huggingface/openai endpoint mapping setup. between stuff like emfasys & fpgas it will be cool to see all of upcoming the inference workload hardware optimizations.

Ignoring FPGA is no longer an option Mark my words - Kintex Ultrascale+ will become the new ICE40 Ultra Plus.

Hope you find success with a niche for this, but "a little bit lighter on FLOPs" is one of the statements of all time.

Impressive! We might have a use for it depending on how fast we can get data in and out of the box. Any info regarding that? My DMs are open!

Hell yes!

Check DM my dude

What's your Moat bro ? Programming FPGA sucks still whereas ASIC reprogram in few micro s

@BrianRoemmele

please be mindful if you process graphics at high speeds :) reality is less hypnotic than a chariot made of LED cascades

Could be fun for radar signal processing. Best of luck and keep us updated!

Some ideas: 1, RFSoC Polyphase Filterbank, FFT, signal categorization,MUSIC, for advanced passive DF for Drone detection, all HDL 2, find a Apache TVM example for quick demo

Isn't FPGA have limited frequency in comparison to other types of general purpose hardware?

@alessandrod wen can also get this instead of good software.

A single Chrome tab will still eat up all that RAM 😂

You should check out @solana.

would you port cuda/cudnn/cublas projects?
