If someone wants to try running a full tpch benchm...
# daft-dev
r
If someone wants to try running a full tpch benchmark, please try the following:
Copy code
# 1. Bring the benchmarking commit into tree.
git cherry-pick 5dc5988168dcac1e549801121630a489244cb9f5

# 2. Install the GitHub CLI tool if you don't already have it.
# The below will install it on macOS/Linux systems with brew.
brew install gh

# 3. Add and commit your changes or whatever it is you want to benchmark.
# You must ensure that the commit you cherry-picked is added and committed to your branch (for the current time being; you can roll this back later).

# 4. Run the workflow.
gh workflow run build-commit-run-tpch.yaml --ref $MY_BRANCH_NAME
This works for me currently on my branch, but want to see if it can extend to other branches as well (i.e., there are no branch-specific assets that I'm depending on that'll fail to resolve when being run elsewhere). An example of a successful workflow looks like: https://github.com/Eventual-Inc/Daft/actions/runs/11906033199 If you scroll down, you'll see some nice outputs.
When this merges into main, you will not need to cherry-pick. You will also be able to run it via the GitHub UI instead of the CLI.
j
If you can run it with
DAFT_ENABLE_RAY_TRACING=1
there are some JSON traces on the head node at
/tmp/ray/logs/session_latest/daft/*
that I would LOVE to get access to after the run finishes
r
I'll need to look into that. I can start playing around with it and throwing them up on the GitHub Actions summary page.
j
One potential trick is that I usually do the download through the dashboard I’m pretty sure then that there is a dashboard endpoint that we can hit to download the file
So maybe we don’t need any shenanigans with scp… just go through the Ray dashboard endpoint
❤️ 1
r
I'll ask Chat to confirm
j
Don’t ask chatgpt, just run Ray yourself
I doubt it will know the answer
Like if you just run
daft up
on daft launcher, you should be able to browse the logs in the dashboard and check the requests your browser is making when looking at the log files I think
image.png
To list the files, it’s sending a GET request to: http://localhost:8265/api/v0/logs?node_id=6cb567218e7a61e6fb6be53766438556dd2b6ce2fe42259c492e9860&glob=daft%2F* which returns a response:
Copy code
{
  "result": true,
  "msg": "",
  "data": {
    "result": {
      "internal": [
        "daft/trace_RayRunner.fa157bd0-3782-4a62-a5c6-e4591506e8ea.2024-11-19T04:56.json"
      ]
    }
  }
}
To list nodes, it’s sending a GET request to:
<http://localhost:8265/nodes?view=summary>
r
@jay The flag
DAFT_ENABLE_RAY_TRACING
is in main, correct?
j
Yes
🚀 1
r
@jay I can pull everything down now. What would you want to include? I can pull all of the
/tmp/ray
dir down and package it up into an artifact if you want??? Or I could just pull down
/tmp/ray/session_*/logs/daft
. Either way works for me.
j
Just the
/tmp/ray/session_*/logs/daft
for now for me
Other folks might want more later on — we should also start digging through that stuff to understand what’s in there that could be useful
And I think as a demo, if you can run TPC-H with this and show us the traces that’d be SICK
👀 1
r
I agree. Additionally, I have been able to display the output CSVs directly on the GitHub Actions UI, so if I'm able to leverage some tooling to view traces directly in the GHA, that would be cool too. Maybe not a prio (and maybe not what you're referring to), but a cool concept nonetheless.
j
The traces require pretty specialized frontends to view it. I’ve been using https://ui.perfetto.dev/
👀 1
r
Ah ya forgot you're writing to the tracing format that Perfetto uses. Nvm