Running the probes during inference maintains the same throughput with no added latency.
That’s enabled by our infra engineering, like kernel-level optimizations and a custom inference server.
Goodfire reports that running its probes during model inference maintains the same throughput with no added latency. The company attributes this to infrastructure engineering, including kernel-level optimizations and a custom inference server.
A reply · the post it answers
Running the probes during inference maintains the same throughput with no added latency.
That’s enabled by our infra engineering, like kernel-level optimizations and a custom inference server.
Source: Goodfire · x.comPublished · added here