Conversation
407202d to
fa556f3
Compare
iamjoemccormick
left a comment
There was a problem hiding this comment.
Dropping a note I finished with my initial pass through this.
a2a3f5e to
4e1d63a
Compare
634f0aa to
cce8e6e
Compare
iamjoemccormick
left a comment
There was a problem hiding this comment.
Submitting review feedback on the first two commits (up to "add xtreemstore provider").
| executeAfter := rst.GetJobExecuteAfter(j.Get()) | ||
| j.Segments = make([]*Segment, 0, len(workRequests)) | ||
| for _, wr := range workRequests { | ||
| if executeAfter != nil { | ||
| wr.SetExecuteAfter(executeAfter) | ||
| } |
There was a problem hiding this comment.
todo: unless I'm missing something doesn't this still do what we were trying to avoid? Generating an timestamp on one node (Remote) that is then passed to the Sync nodes which who's times may not be in sync?
I would prefer instead we make this a DelayExecution field on the job request that uses the protobuf duration type that is propagated to a DelayExecution field on the work requests.
I would also propose this is set by the GenerateWorkRequests method, if appropriate for that RST client type. It doesn't feel right to set this one field in GenerateSubmission().
There was a problem hiding this comment.
I saw this but was distracted by the other comments and forgot to come back to it. Anyway, this required adding a DelayExecution field to both beeremote.JobRequest and flex.WorkRequest which allows the sync node to convert the delay to ExecuteAfter in SubmitWorkRequest. GenerateWorkRequests is still called by remote.
See e85cf75
| zap.Bool("hasWorkResult", entry.WorkResult != nil), | ||
| zap.Bool("hasStatus", status != nil), | ||
| ) | ||
| m.scheduler.AddRescheduleWorkToken(submissionId, time.Time{}) |
There was a problem hiding this comment.
question: so when this happens we still add a token because when the journal is replayed later this WR will still get picked up and presumably the invalid result sent back to Remote?
Edit: I see this was added with the "update job state to running..." commit. Was this a bug? If so it'd be worth a mention in the commit message.
There was a problem hiding this comment.
Yes, the work needs to be processed in order to complete/update remote.
It was a bug. I missed adding m.scheduler.AddRescheduleWorkToken(submissionId, time.Time{}) the first time. Showed up in my testing.
iamjoemccormick
left a comment
There was a problem hiding this comment.
Posting review comments for the remaining commits.
| if _, err := w.beeRemoteClient.UpdateWorkRequest(work.ctx, result.Work); err != nil { | ||
| log.Warn("unable to update remote job status to running; continuing work request without retrying", zap.Error(err)) | ||
| } |
There was a problem hiding this comment.
suggestion(blocking): we discussed this on Slack, but adding a note here so we don't loose track.
Sync intentionally never informed Remote when a work request shifted to "running" to reduce load on Remote and the jobs DB. I would prefer we keep it that way, as for the most part Remote never needs to know about this unless a user checks the job status.
What we could do is implement this issue https://github.com/ThinkParQ/bee-remote/issues/14 so if the user runs remote status or remote job list Remote will reach out and refresh the job and work request statuses. The GetJobsRequest message already has a UpdateWorkResults field that was intended to control if only the latest results known to Remote are returned, or if it also probes the Sync nodes.
We could either set UpdateWorkResults by default for CTL commands, or add a new --refresh-results flag for this. My thinking is for the most part users only care to know once jobs reach a terminal state, the intermediate states aren't that interesting unless you're trying to debug specific job issues.
cce8e6e to
b01a2c0
Compare
2ba6ab2 to
9d97563
Compare
iamjoemccormick
left a comment
There was a problem hiding this comment.
Just a handful of remaining observations, my comment on GenerateSubmission() is the main blocker.
| } | ||
|
|
||
| } else { | ||
| workRequests = rst.RecreateWorkRequests(j.Get(), j.GetSegments()) |
There was a problem hiding this comment.
todo: the RecreateWorkRequests() path was updated to handle SetDelayExecution(), but I think we also need to set this in the GenerateWorkRequests() path?
I'm actually not sure if the RecreateWorkRequests() path in GenerateSubmission() is ever executed anymore, elsewhere just calls RecreateWorkRequests() directly.
| if client.IncludeInBulkRequest(walkCtx, jobRequest) { | ||
| rstId := jobRequest.GetRemoteStorageTarget() | ||
|
|
||
| bulkRequestStatesMu.Lock() |
There was a problem hiding this comment.
suggestion(non-blocking): to avoid lock contention in a hot path consider making this a RWMutex as after the initial population of bulkRequestStates the map is only read.
Assisted-by: Claude:claude-opus-4-6
|
|
||
| type SubmitBulkRequestFn func(ctx context.Context) | ||
| type EmitBulkRequestFn func(ctx context.Context, request *beeremote.JobRequest) | ||
| type AppendBulkRequestFn func(ctx context.Context, request *beeremote.JobRequest) |
There was a problem hiding this comment.
suggestion(non-blocking): we should mention AppendBulkRequestsFn is called under a lock so it should be fast. For example of a provider did I/O in append it could really slow everything down.
Assisted-by: Claude:claude-opus-4-6
8c941da to
d23cac0
Compare
| return | ||
| } | ||
|
|
||
| // getBuilderResults generates a status message based on the builder submission counters. |
There was a problem hiding this comment.
Here were some generated examples,
| Builder / SchedulerResult | State | Status message |
|---|---|---|
submitted=250clean, mid-walkReschedule=true |
RESCHEDULED |
waiting for builder job to continue; 250 job request(s) submitted |
submitted=100, errors=5Reschedule=true |
RESCHEDULED |
waiting for builder job to continue; builder cancelled: 100 job request(s) submitted; 5 submitted with errors |
submitted=42clean, abortedErr=MarkBuilderCancelled("job builder request was aborted: context canceled") |
CANCELLED |
job builder failed to complete: job builder request was aborted: context canceled; 42 job request(s) submitted |
submitted=10, errors=3Err=MarkBuilderFailed("job builder request was aborted: unable to abort bulk operation: connection refused") |
FAILED |
job builder failed to complete: job builder request was aborted: unable to abort bulk operation: connection refused; builder cancelled: 10 job request(s) submitted; 3 submitted with errors |
submitted=7, alreadyExist=2Err="unexpected nil bulk result"unclassified |
FAILED |
job builder returned unclassified error: unexpected nil bulk result; 7 job request(s) submitted; 2 already exist |
submitted=180, alreadyComplete=6, notAllowed=2, errors=4Err=nil, Reschedule=false |
CANCELLED |
completed with errors; builder cancelled: 180 job request(s) submitted; 6 already complete; 2 not allowed; 4 submitted with errors |
submitted=500clean runErr=nil, Reschedule=false |
COMPLETED |
completed successfully; 500 job request(s) submitted |
alreadyComplete=1000nothing newErr=nil, Reschedule=false |
COMPLETED |
completed successfully; 1000 already complete |
notAllowed=50all conflictedErr=nil, Reschedule=false |
CANCELLED |
completed with errors; builder cancelled: 50 not allowed |
submitted=999, alreadyOffloaded=1Err=nil, Reschedule=false |
COMPLETED |
completed successfully; 999 job request(s) submitted; 1 already offloaded |
0 matcheslocal path /data/empty-dirmid-walkReschedule=true |
RESCHEDULED |
waiting for builder job to continue; builder cancelled: no matches found in local path: /data/empty-dir |
0 matchesremote path s3://bucket/incoming/Err=nil, Reschedule=false |
CANCELLED |
completed with errors; builder cancelled: no matches found in remote path: s3://bucket/incoming/ |
0 matchesdownload fallback to local walk /local/mirrorErr=nil, Reschedule=false |
CANCELLED |
completed with errors; builder cancelled: walked local path since --remote-path was not provided; No matches found in path: /local/mirror |
0 jobs countedaborted immediatelylocal path /data/xErr=MarkBuilderCancelled("job builder request was aborted: context canceled") |
CANCELLED |
job builder failed to complete: job builder request was aborted: context canceled; builder cancelled: no matches found in local path: /data/x |
d23cac0 to
6d4e8a0
Compare
52b9071 to
4bf837e
Compare
4c5cab0 to
394073f
Compare
…iginal request The ctl previously set the remote path to the in-mount path for upload requests. This prevented the s3Client.GenerateWorkRequests from having a chance to update the remote path based on the last job or normalizing the key.
BuildJobRequest set RemoteSize, RemoteMtime and IsArchived unconditionally so a missing remote object left lockedInfo.RemoteMtime holding a zero timestamp rather than nil. Both IsFileAlreadySynced and HasRemotePathInfo key off that distinction.
The work journal entries were passed by value instead of reference so a copy of ExecuteAfter was modified instead of the database. Few cleanups with no behavior changes * Added appendMessage helper to help aggregate error messages for ctl.
Walks were previously limited to file or object count and then a resume token would be returned. This prevents the caller from terminating as needed. So now, * GetWalk returns an explicit stop function. * All emitted results contain a valid resume token.
Add a trailing slash when comparing a directory with the resume token to avoid skipping directories because of situations like a directory b sitting next to a file b.txt. Without the trailing slash, the directory is skipped on resume even though it is emitted after that file.
GetLockedInfo previously handled rstId logic on a provided config which is removed so the caller can explicitly handle it. * Rename GetLockedInfo as GetPathState to better reflect the intent. * Return PathState object with complete entryInfo so caller can use it.
Providers previously could not distinguish between a context error for cancelling a job and sync shutting down. This made it impossible to shutdown gracefully. * Add shutdown context explicitly to ExecuteWorkRequestPart and IsWorkRequestReady. * Update sync worker to support gracefully shutdown and resume on restart.
Previously the local file state was immediately changed in preparation for work requests. So if the caller needed to undo the changes there was no clear path. * Extract the required work request changes and steps to undo into returned functions. * Leave existing file contents intact until data is written when downloading. This is more efficient than resizing and then truncating the file, and improves cleanup when the job is cancelled. * IsFileOffloaded now requires the file to exist.
* Classify all job submission responses as rst sentinels. * Wrap remote connection issues with ErrUnavailable.
* Track average sync node work load saturation across a short and long time window. * Pass load information to sync workers and builder jobs for rescheduling logic.
…elay contexts Gracefully shutting down sync nodes should allow any inflight tasks to complete before stopping. This includes rpc and api calls that may be slow. WithCancellationDelay and grace checkpoints will be added to critical paths so they can complete. * Add reporting to logs to show remaining inflight tasks.
… work failures Clients need have to opportunity to handle work failures. In many instances the failures can be automatically be set into a valid state such as in uploads or when the local file state can be restored. This change will be especially important for adding bulk operation support since requests will potentially hold critical resources that should always be released regardless of the individual file state. Note the client can still fail the request when the file state requires manual intervention; only that the client is making that determination instead of remote enforcing it. * Add support for rst clients to remedy work request failures instead of assuming they should be failed. * Implement started flag to work request part to indicate when work has commenced. * Improve S3Client cleanup when completing a job request by, * replacing an incomplete download with a stub file. * restoring the original stub if there was one. * removing the file if it didn't exist. * otherwise restore the file's size and mtime. * Fix remote's work assignment process so it stops assigning work immediately when a request fails to be assigned.
…equests with failed-precondition
* Add support types and interface for bulk operations. * Update s3 client to satisfy Provider changes. * Update mock client to satisfy Provider changes with tests.
Split builder.go into supporting files with clear responsibilities: * buildercontroller.go coordinates the walk, the worker pool and the bulk operations for one builder job, and collects the results. * builderjobrequest.go turns a walked path into a job request, including resolving path state and planning the local file state changes each request needs, with an undo for every step so a failure walks them back. * builderbulk.go holds the bulk operation registry and manager, persisting each operation's state under the mount point so it survives a restart. * Job requests are built and submitted by the goroutine that prepared it, so the submission outcome is known while everything needed to roll that path back is still in scope. * Sync worker and job builder can now distinguish between work cancelled and sync shutting down. This allows work to conclude before the shutdown without stepping backwards.
…AXPROCS The AWS SDK's default MaxIdleConns/MaxIdleConnsPerHost caps connection reuse well below what a single RST config's concurrency needs, since it only ever talks to one host. Size it off the same multiplier the request build controller uses for its own concurrency so the two stay in sync.
* Add support for clients to reschedule work requests.
* Update s3 client so work is rescheduled on graceful shutdown instead of failing.
* Uploads resume after the last completed work request part.
* Downloads resume after the last byte written.
Previous downloading empty an object would fail to overwrite the size and mtime of the local file. * Add support for downloading empty files to PlanFileStateForWorkRequests.
Add support for configuration files to reference enum values by name. The enum name is case insensitive and hyphens can be used inplace of underscores.
Xtreemstore is s3-compatible and therefore extends the s3 client implementation.
* Add xtreemstore provider
* Add bulk retrieval support for downloading archived objects.
* Independently manage multiple asynchronous bulk restore operations across multiple builder jobs.
* Active bulk operations will reschedule until the last object is downloaded.
* Bulk operation managers that are blocked by another's active retrieve session will reschedule
for a later time.
* Only one retrieve session can be active at a time so the bulk operation manager keeps the
builder job alive until all archived objects are downloaded. This mechanism is necessary to
manage xtreemstore's tape buffer.
* State and request statuses are persistently stored on the beegfs mount so the bulk
operation can be restored after a shutdown or crash.
Inflight work requests must be given a chance to complete while its sync worker is shutting down. So, * Only a single attempt to update a job request's work status is permitted during shutdown. * Any failures are logged.
…testing. This race condition was exposed when require.Eventually replaced the explicity sleep.
* Fix the name the interface a new RST type must implement. * Update GenerateWorkRequests description to include availableWorkers.
cb8d2be to
1a5a348
Compare
What does this PR do / why do we need it?
Required for all PRs.
Related Issue(s)
Required when applicable.
Where should the reviewer(s) start reviewing this?
Only required for larger PRs when this may not be immediately obvious.
Are there any specific topics we should discuss before merging?
Not required.
What are the next steps after this PR?
Not required.
Checklist before merging:
Required for all PRs.
When creating a PR these are items to keep in mind that cannot be checked by GitHub actions:
For more details refer to the Go coding standards and the pull request process.