https://www.getdaft.io logo
Join Slack
Powered by
# daft-dev
  • a

    Abhijeet Patil

    03/24/2026, 5:45 PM
    Hi Team, I am running ETL workload using daft distributed mode on ray cluster. The ETL jobs picks up modified records in last 24 hours from mysql source, applies minor transforms like type casting, adding partition column (date_part), and finally does delta merge on primary keys. Since delta merge is not natively supported in daft, I am using the data sink workaround (https://github.com/Eventual-Inc/Daft/issues/4047#issuecomment-3094861350). I am seeing ~5% data loss in columns only with datatype DECIMAL(38, 20). I typecast it to
    daft.DataType.decimal128(38, 20)
    after reading from mysql. If distributed delta merge using deltalake fails on a worker due to delta contention, I am doing retries (I am not using dynamodb as lock provider). Is there a known issue with datatype decimal(38, 20) and how to avoid it? Please let me know if additional details are required.
    r
    j
    • 3
    • 16
  • a

    Abner Ayala

    03/25/2026, 9:37 PM
    Hi Guys I'm still struggling a lot to understand the daft.cls API for model inference if I want to avoid passing hard-coded values. Below I have a simple dummy resnet model that expects as input a batch of images and returns 2 outputs a json string with the probs and the embeddings. More importantly I want to control max_concurrency and batch size during runtime. That means I need to create the class without decorators and then wrap it during runtime. Questions: 1. I know I need to use daft.cls and daft.method.batch but how to do this during runtime when not able to use decorators which expect import time variables; or how to use the decorators but overwrite some of these parameters like max_concurrency and batch_size at runtime. 2. How to mix unested=True when using daft.method.batch which expects output to be a series or a list but unnest-True expects output to be a Struct?
    Copy code
    class SimpleModel:
        def __init__(self, model_checkpoint_uri: str):
            self.device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
            self.model = load_model(model_uri).to(self.device)
            self.model.eval()
    
        def predict(self, images):
            if len(images) == 0:
                return []
    
            images_tensor = torch.from_numpy(np.array(images.to_pylist())).to(self.device)
            batch_size = images_tensor.shape[0]
            with torch.inference_mode():
                probabilities, features = self.model(images_tensor)
                features = features.half().detach().cpu().numpy().tolist()
                rounded_probabilities = probabilities.round(decimals=4).detach().cpu().numpy().tolist()
                predicted_class_id = probabilities.argmax(dim=1).detach().cpu().tolist()
                predicted_class_name = [CLASSES[p] for p in predicted_class_id]
    
                predictions = []
                for i in range(batch_size):
                    prediction = json.dumps({
                        "class_id": predicted_class_id[i],
                        "class_name": predicted_class_name[i],
                        "probability": rounded_probabilities[i],
                    })
                    predictions.append(prediction)
    
                return predictions, features
    During Runtime: Not working
    Copy code
    ModelClass = daft.cls(EyesClosedClassifier, gpus=1, max_concurrency=cfg.num_gpus)
        model = ModelClass(model_checkpoint_uri=cfg.pipeline.model.checkpoint_uri)
        model_inference = daft.method.batch(
            model.predict,
            return_dtype=daft.DataType.struct({
                "predictions": daft.DataType.string(),
                "features": daft.DataType.list(daft.DataType.float32()),
            }),
            unnest=True,
            batch_size=cfg.batch_size,
        )
    r
    c
    e
    • 4
    • 19
  • a

    Abner Ayala

    03/30/2026, 8:45 PM
    Hi Team Another bug 😕. I have a daft.DF that has several lightweight df.func , no GPUs used at all. But I never materialize it. Then I have another daft.DataFrame that I want to use to join and add a single column to the original pipeline. So df1.join(df2, on=url), but getting error. Seems I need to materialize df1 before join? but df1 has 100M+ rows, ideally I don't want to materialized till df.write?
    PanicException: Tried to get unmaterialized stats
    👀 1
    d
    e
    r
    • 4
    • 6
  • r

    Ruthvik

    04/06/2026, 9:42 AM
    Hi team! I was thinking if we should add tests to 6606 and in general to Write stats/metrics. Any ideas on how to do that?
    d
    • 2
    • 3
  • g

    Garrett Weaver

    04/27/2026, 8:37 PM
    👋 looking to add Jaccard similarity and I was curious if it would be preferred to "bake in" generating work/n-grams with some configurable params (similar to https://www.jpt.sh/projects/jellyfish/functions/#jaccard-similarity) or expect users to handle themselves and pass in two columns with type
    list(str)
    ? I somewhat prefer the explicit handling by users over hiding some details, but curious on thoughts.
    c
    • 2
    • 1
  • m

    Mehul Batra

    05/06/2026, 8:33 AM
    Hi Team, Mehul from DigitalOcean's . We're in the middle of evaluating Daft & Raydata for a multimodal ingestion & transformation on Kubernetes (Heterogeneous GPU/CPU, Batch Inference, sinking to Lance + Iceberg with Gravitino catalog registration). Daft is technically the better fit on most dimensions that matter to us, RayData seems to be good too but need alot of wrapper on low level API, but surely winning in corporate marketing with anyscale summit. Before we commit, I'm trying to get a clearer picture on a few things, and the docs only go so far: 1. Real-world production deployments on self-hosted K8s The public reference set (Amazon BDT, CloudKitchens, Together AI, Essential AI) is well-documented but mostly 12–18 months old. Are there newer teams running Daft on self-hosted Kubernetes not Daft Cloud, not Anyscale at meaningful scale? On that note: we're aware that ByteDance's LAS (Lake for AI Service) team has been one of the most active contributors on the Daft↔️Lance integration surface distributed FTS, compaction, schema evolution, limit pushdown, and most recently
    GravitinoGvfs
    local filesystem writes in v0.7.10 (
    @qingfeng-occ
    ). Their pattern Lance as storage, Ray as runtime, Daft hardening the client layer is close to our target architecture. If anyone from that team or others running a similar stack is in this Slack, I'd love to connect. 2. Flotilla vs set_runner_ray() — current recommendation for DOKS We're open-source-only. Is the Flotilla native K8s path production-ready today for a team that can't fall back to Daft Cloud? Or is the practical answer still
    set_runner_ray()
    + KubeRay for anything serious? Any known stability cliffs under sustained load? 3. Checkpointing roadmap for Lance/Iceberg sinks
    daft-checkpoint
    scaffolding landed in v0.7.6 — great signal. Our write path is Lance + Iceberg + OpenSearch, and the current Parquet-only limitation means we'd still need external idempotency for v1 regardless. Is Lance/Iceberg write support on the near-term checkpoint roadmap, or is that explicitly out of scope for now? We'd plan to contribute back when things fall in place. Happy to share our UC matrix and PoC findings once we're further along.
    👋 1
    e
    r
    c
    • 4
    • 7
  • d

    Desmond Cheong

    05/14/2026, 4:32 PM
    @Robert Howell could you help take a look at this PR? https://github.com/Eventual-Inc/Daft/pull/6760 It adds support for our native extensions on the distributed runner
    👋 6
    👀 2
    ❤️ 1
    r
    • 2
    • 1
  • u

    北南

    06/04/2026, 5:01 PM
    https://github.com/daft-engine/daft-lance/pull/18 @Robert Howell could you help take a look at this this PR? thanks!
    ✅ 1
  • u

    北南

    06/05/2026, 4:59 AM
    https://github.com/daft-engine/daft-lance/pull/21 @Robert Howell Support Distributed Segmented BTREE Index Building
    ✅ 1
  • u

    北南

    06/05/2026, 9:54 PM
    https://github.com/daft-engine/daft-lance/pull/23 @Robert Howell Integrates Lance's log-structured ingestion framework (MemWAL) with
    daft-lance
    to enable high-throughput parallel writes. thanks a lot!!
    ✅ 1
    🙌 1
    r
    j
    • 3
    • 8
  • a

    Abhijeet Patil

    06/09/2026, 5:23 PM
    Hi Team, is there a drop-in replacement of spark streaming in daft?
    ➕ 1
    u
    j
    • 3
    • 4
  • u

    北南

    06/13/2026, 5:10 AM
    Need help for code review: https://github.com/daft-engine/daft-lance/pull/24 https://github.com/daft-engine/daft-lance/pull/25 https://github.com/daft-engine/daft-lance/pull/26
    j
    • 2
    • 5
  • a

    Abhijeet Patil

    07/02/2026, 6:12 PM
    Hi Team, can someone please help with code review: github.com/Eventual-Inc/Daft/pull/7187
    s
    • 2
    • 5
  • a

    Abhijeet Patil

    07/10/2026, 5:47 PM
    Hi Team, do we have functions in daft to detect PII?
    s
    • 2
    • 3
  • a

    Abhijeet Patil

    07/13/2026, 10:13 AM
    Hello Team, we are writing an engineering blog about the gains we saw after migrating from spark to daft at MakeMyTrip. We would like to submit it to daft's website's blog section. (eventual.ai/blog) Can someone please tell us steps to do that and guidelines for the same?
    • 1
    • 1
  • s

    Satyendra Kumar

    07/13/2026, 10:21 AM
    Hi Team, please take a look at these: github.com/Eventual-Inc/Daft/discussions/7259 github.com/Eventual-Inc/Daft/issues/7260
  • a

    Abhijeet Patil

    07/17/2026, 3:33 PM
    Folks with experience running Daft on GCP, which family of compute engine offers best price to performance ratio (intel, amd, axion, etc)? Asking for ETL pipelines.
  • d

    dsohong

    09/02/2026, 3:34 AM
    Hi Team, I need help with a code review for Daft's Iceberg read/write features. github.com/Eventual-Inc/Daft/pull/7419 github.com/Eventual-Inc/Daft/pull/7421 github.com/Eventual-Inc/Daft/pull/7429 thx
    👍 2
  • p

    peter-tang

    09/08/2026, 2:18 PM
    Hi Daft team — my PR #7439 (push  Limit  down to the row-preserving side of Left/Right joins, addressing #2275) is green and mergeable; could someone take a look when you have a chance? https://github.com/Eventual-Inc/Daft/pull/7439:arrow_upper_right:
    ✅ 1
    c
    • 2
    • 1
  • d

    dsohong

    09/09/2026, 3:43 AM
    Hi all — no urgency at all, but if anyone has a spare cycle I'd really appreciate a look at two small Iceberg fixes: - #7421 — count_rows() on an Iceberg table silently ignored a filter on an identity-partitioned column and returned the whole table's row count. - #7429 — after a partition spec evolved to add a field, a filter on that field returned rows that don't actually match it. github.com/Eventual-Inc/Daft/pull/7421 github.com/Eventual-Inc/Daft/pull/7429 Happy to walk through the reasoning on either one, or to split them further if that makes review easier.
  • p

    peter-tang

    09/09/2026, 5:41 AM
    Hi team — could I get a re-review on PR #7398 (fix: keep pushed filter columns in projection pushdown)? https://github.com/Eventual-Inc/Daft/pull/7398↗ It fixes #6757:  read_parquet(...).where(col("b") < 1).select("a")  used to fail because projection pushdown stopped reading the filtered column. @rchowell both your earlier asks are addressed — the greptile JSON concern is fully reverted, and the count-pushdown CI failure is fixed. All checks green, mergeable with no conflicts, well-tested (unit + fixed-point race + JSON/Parquet e2e). Been quiet since Aug 24 — would appreciate a look when someone has a moment. Thanks!
  • p

    peter-tang

    09/09/2026, 7:27 AM
    Hi team — I've opened #7489, a small two-line fix for the  integration-test-catalogs  failures that have been affecting PRs since ~09-06 (tracked in #7476). The cause is infrastructure rather than anyone's code: the Iceberg test image builds  FROM python:3.10-bullseye , whose apt suite has gone stale (openjdk-11 packages 404), so  docker compose up  aborts and the catalog tests fail. The fix moves it to  python:3.10-bookworm  and CI is fully green, including that job — would anyone have a moment to review and merge it? Thank you!
  • f

    fanng

    09/10/2026, 8:25 AM
    Hi team, We’re building a multimodal processing pipeline with Daft + Lance. The following PRs add support for Lance table update and insert_overwrite semantics in Daft. Appreciate your review. github.com/daft-engine/daft-lance/pull/58 github.com/daft-engine/daft-lance/pull/62
    ✅ 1
    🙌 1
  • f

    fanng

    09/10/2026, 8:32 AM
    Daft-lance was maintaining a second implementation of Lance’s merge-column write path. This exposed too many Lance internals, making it difficult to maintain and extend, and introduced known bugs such as #59 and #20. github.com/daft-engine/daft-lance/pull/60 proposes using Lance’s existing merge_columns API to implement a fast path for adding columns, Please help to review, thx!
    ✅ 1
  • p

    peter-tang

    09/12/2026, 1:48 PM
    Hi Team — I added the e2e test you asked for on #7441 (prefix  starts_with  pruning via min/max stats). Would you mind taking another look when you have a moment? Thanks! https://github.com/Eventual-Inc/Daft/pull/7441:arrow_upper_right:
  • f

    fanng

    09/14/2026, 2:23 AM
    Hi! Is there a plan for a new daft-lance release? 0.4.0 went out on June 5, and main has 11 unreleased commits since — Lance Namespace support (#35), atomic insert_overwrite (#62), segmented bitmap indexes (#31), plus several merge fast-path correctness fixes (#44, #45, #46, #56, #60). Since daft's lance extra tracks the latest daft-lance, none of this reaches daft users until a release is cut. If maintainers are short on time, I'm happy to help — I can open the release PR (version bump + changelog) so all that's left is publishing the GitHub Release. Just ping me for whatever else is needed.
    r
    • 2
    • 10
  • f

    fanng

    09/18/2026, 1:09 AM
    Hi team, could you help review update lance column PR? Typically, it's used to re-running a model and writing fresh embeddings or scores back to the same column, backfilling or correcting a column at scale, it's import in our Daft&Lance pipeline, thanks! github.com/daft-engine/daft-lance/pull/58
  • j

    Jianyao Ma

    09/22/2026, 1:43 PM
    Hi team, I reproduced #7458 locally and have a prototype fix. It keeps one writer per input while allowing bounded concurrency across inputs. Details are in my comment. Could someone familiar with BlockingSink take a look and let me know if this approach makes sense? Thanks!
  • f

    fanng

    09/23/2026, 7:29 AM
    Hi team,
    daft-lts
    release is not updated for the limit of pypi storage limit, could any one help to handle this? thanks. github.com/Eventual-Inc/Daft/issues/7546
    s
    • 2
    • 5
  • j

    Jianyao Ma

    09/24/2026, 6:14 AM
    Quick follow-up on #7458 — the key invariant in the prototype is that each InputId still has exactly one writer, while concurrency is bounded across different inputs rather than within one input. If that matches the intended semantics for BlockingSink, I can clean this up and open a PR.
    b
    • 2
    • 2