Skip to content
Discussion options

You must be logged in to vote

@NV-sschneider can answer this better, but we are having some discussions about exactly what direction we want to support here in terms of other styles of upstream resources (ex. k8s clusters).
On the idea of routing:

  • We do plan to keep the boundary of choosing the right upstream and being a passthrough.
  • We plan to update routing to have a handful more options rather than the naive version right now. We are working on these changes right now, but the goal is eventually to support things like kv cache, throughput measurement (effectively how strong is a given machine), queue depth, and ideally measure how large a workload might be and do some prediction, cost, context size, throughput etc…

Replies: 1 comment 2 replies

Comment options

You must be logged in to vote
2 replies
@NV-sschneider
Comment options

@vonargo
Comment options

Answer selected by vonargo
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
3 participants