<@U042126MG49> <@U041QSEF2H2> do you have any thou...
# daft-dev
k
@jay @Sammy Sidhu do you have any thoughts about populating io config keys with the relevant environment variables upon construction time instead of at usage time? e.g. if you call
azure = <http://daft.io|daft.io>.AzureConfig()
without a
storage_account
but
AZURE_STORAGE_ACCOUNT_NAME
exists in the env, it will set
azure.storage_account
to the current value of the env var
I think a potential downside would be that if a user has a workflow were they change env var values between the construction of the io config and it usage, that might break their workflow. However I think it could make the behavior of the io config values more clear since the user knows exactly what is set
j
The reason it’s done at usage time is because of distributed environments where you might want to pick it up on the cluster instead of on the client I think the current behavior is: • nothing is specified: pick it up on cluster • Something is specified: user does their own credentials flow and provides us credentials on the client-side I think we don’t always adhere to this now though, because we do some client-side IO for delta and iceberg. In the past we took great care to perform even globbing in the cluster (see: RayRunnerIO)
k
Ah I see