Skip to content

systemd: prevent hanging default target - #3874

Open
LuanVSO wants to merge 3 commits into
abraunegg:masterfrom
LuanVSO:master
Open

LuanVSO wants to merge 3 commits into
abraunegg:masterfrom
LuanVSO:master

Conversation

@LuanVSO

@LuanVSO LuanVSO commented Sep 14, 2026 •

Copy link
Copy Markdown

fix some ordering/ dependency issues with the systemd service, which could cause delayed start of the desktop session.

@abraunegg

Copy link
Copy Markdown
Owner

@LuanVSO

Can you please provide some deeper context here, and what platform this change has been tested on

@abraunegg

abraunegg commented Sep 14, 2026 •

Copy link
Copy Markdown
Owner

@LuanVSO

Further to the above question, I have now gone considerably deeper into the implications of replacing:

ExecStartPre=/bin/sh -c 'sleep 15'

with:

After=dbus.socket

across both supplied systemd unit files.

There are a few concerns that I think need to be addressed before this can be merged.

The first point is that this does not appear to be a systemd version compatibility problem. After= is supported by systemd versions far older than anything currently relevant to the supported Linux distributions.

The concern is instead around systemd dependency semantics, service-manager scope, and whether dbus.socket actually represents the D-Bus instance that the OneDrive client needs.

1. After=dbus.socket is only an ordering dependency

After= does not cause another unit to be started.

In other words:

After=dbus.socket

means:

If dbus.socket is part of the current systemd transaction, order this service after it.

It does not mean:

Start dbus.socket, wait until it is available, and then start OneDrive.

If dbus.socket is not installed, not enabled, or otherwise not part of the transaction, After= alone does not make it a runtime requirement.

This is particularly important because removing the existing 15-second delay means the OneDrive process may now start immediately in those circumstances.

If the intention is that systemd should actively bring the D-Bus socket into the transaction, then I would expect this to require consideration of something such as:

Wants=dbus.socket
After=dbus.socket

rather than After= by itself.

I would not want to make D-Bus a hard Requires= dependency, because D-Bus / GUI notification availability should not prevent the client from synchronising.

2. The two OneDrive service files run in different systemd manager scopes

This is probably the most important issue.

Normally:

onedrive.service

is installed as a user systemd unit, whereas:

onedrive@.service

is installed as a system systemd unit.

This means the same line:

After=dbus.socket

does not necessarily refer to the same thing.

For the user service:

systemctl --user ...

dbus.socket refers to the D-Bus socket known to the user systemd manager.

That is potentially appropriate because the client uses the session/user D-Bus environment for desktop notification support.

For the templated system service:

systemctl ...

dbus.socket refers to the socket known to the system systemd manager.

That is the system D-Bus bus.

The OneDrive client, however, tests the user/session D-Bus environment using values such as:

XDG_RUNTIME_DIR
DBUS_SESSION_BUS_ADDRESS

and GUI notification availability depends on that user/session bus being usable.

Therefore I do not currently understand how:

After=dbus.socket

inside onedrive@.service

is equivalent to the existing startup delay.

It appears to order the OneDrive system service after the system bus, while the historical delay was introduced to allow D-Bus to become usable for the client environment.

Those are not necessarily the same thing.

3. RHEL-family behaviour needs particular consideration

There is also existing special-case behaviour in the OneDrive build/install process for RHEL-family systems.

Historically there have been differences around systemd user unit locations, and the build logic contains specific handling for RHEL/CentOS-family installations.

This makes it particularly important that this change is tested not just using:

systemctl --user start onedrive

on a modern desktop distribution, but also using the system-level templated service:

systemctl start onedrive@<user>.service

On a system-level service, dbus.socket is the system bus socket and does not inherently guarantee that the user's session bus at something equivalent to:

/run/user/<uid>/bus

exists or is ready.

4. dbus.socket is not guaranteed to exist in every supported configuration

There are also supported Linux configurations where systemd is present but the D-Bus session implementation/package layout may differ.

For example, Debian-family systems can use the systemd-integrated dbus-user-session model, but alternative D-Bus session configurations also exist.

Other supported platforms may use:

  • OpenRC;
  • BSD rc systems;
  • SysV-style init;
  • systemd as an optional rather than mandatory init system.

Those systems are not necessarily directly affected by these files, but it reinforces why the supplied systemd units cannot assume that every supported installation has identical D-Bus unit topology.

The relevant question is therefore not simply:

Does this distribution use systemd?

It is:

Does this service execute under the user or system manager, does that manager contain a unit called dbus.socket, and is that socket actually the D-Bus instance the OneDrive client subsequently uses?

5. Behaviour when the referenced socket does not exist

My understanding of the proposed unit is that if the applicable systemd manager does not have dbus.socket, After=dbus.socket alone does not cause the OneDrive service to fail simply because the referenced unit is absent.

Instead, there is effectively nothing useful for OneDrive to wait for.

The service can therefore start immediately.

At that point the application performs its own D-Bus availability check. If the session D-Bus environment is not ready, the client continues operating but GUI notifications are disabled.

That creates a potential regression compared with the existing delay:

boot/login
  |
  +-- OneDrive starts
  |
  +-- user D-Bus not ready yet
  |
  +-- OneDrive detects D-Bus unavailable
  |
  +-- GUI notifications disabled
  |
  +-- D-Bus becomes available shortly afterwards

The synchronisation process itself may continue perfectly normally, so this could be extremely easy to miss during testing.

The visible symptom could simply be that GUI notifications no longer work after boot until OneDrive is restarted.

6. Please clarify how the system service case is expected to work

The main technical question I would like answered is:

onedrive.service normally executes under the user systemd manager, while onedrive@.service executes under the system systemd manager. Therefore After=dbus.socket in these two units refers to two different D-Bus sockets. How does ordering onedrive@.service after the system dbus.socket guarantee readiness of the user/session D-Bus that OneDrive accesses through XDG_RUNTIME_DIR / DBUS_SESSION_BUS_ADDRESS?

I think this needs to be clearly explained before we can treat the two changes as equivalent.

7. Testing I think is required

In addition to the testing details I have already requested, I think this should specifically be exercised using both supplied service models:

systemctl --user enable --now onedrive

and:

systemctl enable --now onedrive@<user>.service

Testing should be performed from an actual boot/login condition rather than simply restarting the service once the desktop session and D-Bus environment are already established.

For each test I would want to know the output of:

systemd --version

and, for the user service:

systemctl --user status dbus.socket
systemctl --user cat dbus.socket
systemctl --user show -p ActiveState,SubState dbus.socket

and for the system service:

systemctl status dbus.socket
systemctl cat dbus.socket
systemctl show -p ActiveState,SubState dbus.socket

It would also be useful to confirm from the OneDrive runtime environment that:

XDG_RUNTIME_DIR
DBUS_SESSION_BUS_ADDRESS

are available and point at the expected user-session bus.

The critical validation is not simply whether:

onedrive.service

shows as active (running).

It is whether the same D-Bus functionality that the existing 15-second delay was intended to protect is reliably available immediately after boot/login.

8. Possible direction

For the user service, something along the lines of:

Wants=dbus.socket
After=dbus.socket

may be a more complete systemd relationship than After= alone, assuming dbus.socket is confirmed to be the appropriate session socket on the supported target systems.

I am much less convinced that the same relationship belongs in:

onedrive@.service

because that service executes in the system manager and therefore resolves dbus.socket in a different scope.

There may also ultimately be a better application-level solution here.

The current 15-second sleep is fundamentally a workaround for a startup race. A more robust design would be for the client to tolerate the user/session D-Bus not being immediately available and re-evaluate notification availability later, rather than permanently disabling notifications because the bus was unavailable during the initial startup check.

That would remove the arbitrary delay without relying on distribution-specific D-Bus/systemd unit topology.

One additional piece of context is important here.

Having reviewed why the existing 15-second ExecStartPre exists, I do not think it should be considered solely a workaround for waiting for D-Bus.

D-Bus availability was one of the original symptoms that led to the delay being introduced, but the delay provides a broader and intentional startup grace period.

In practice it allows time for:

  • systemd startup activity to settle;
  • the system and/or user D-Bus environment to become operational;
  • network connectivity to become genuinely usable;
  • the display manager and graphical user session to initialise;
  • XDG_RUNTIME_DIR / DBUS_SESSION_BUS_ADDRESS and other session state to become available;
  • other boot/login services to complete their startup work;
  • boot-time CPU and filesystem I/O contention to reduce.

Only after that grace period does the OneDrive monitor start its own initialisation.

This distinction is important because:

After=dbus.socket

does not provide equivalent behaviour.

It provides an ordering relationship against one particular systemd unit, in one particular systemd manager scope.

Even if that ordering is correct for D-Bus on a particular distribution, it says nothing about the readiness of the network, graphical session, session environment, other user services, or general boot/login activity.

There is also a notification correctness consideration.

If OneDrive starts before the user's D-Bus notification environment is available, the application can detect that GUI notifications are unavailable and continue synchronising. However, simply detecting D-Bus becoming available later does not solve the problem cleanly: notification-worthy events which occurred before D-Bus became usable have already happened.

For that reason I do not think adding application logic later to notice that D-Bus has appeared is equivalent to starting the application once the surrounding environment has had a chance to become ready.

If the objective of this PR is specifically to prevent the existing ExecStartPre from contributing to the time taken for default.target to become active, I think that should be treated as a separate problem.

The requirement should be:

Preserve the intentional startup grace period, while avoiding unnecessary blocking of the surrounding systemd target.

For example, a delayed/timer-based activation model may be worth investigating if there is a demonstrated issue with the current arrangement.

However, replacing the existing 15-second startup grace period solely with:

After=dbus.socket

does not preserve the behaviour that is currently being provided, and therefore I do not think the proposed change is equivalent as it currently stands.

For this PR specifically, however, I think we first need to establish that the proposed replacement is semantically equivalent to what it removes across both supplied systemd service models.

@LuanVSO

LuanVSO commented Sep 14, 2026

Copy link
Copy Markdown
Author

i tested this on fedora and debian the current solution was delaying plasma start up by 15s:

systemd-analize --user plot onedrive-delay
  1. After= is only an ordering dependency
    yes that was intentional, as dbus is not a hard dependency, but i will add it since you've demonstrate concern.

  2. Different systemd manager scopes
    then the current solution also doesn't work, because the system unit will be started at boot, if the user does not login within the 15 seconds timeout, the user dbus session will not exist.

  3. dbus.socket is not guaranteed to exist everywhere
    i could add basic.target to the After statement, but debian does have both dbus.socket and dbus.service.
    the sysvint, openrc and bsd rc targets are not affected by this change.

both units have an after statement for network-online so that already guarantees that the network is in a functional state.

but i agree that the template ondrive@ change is not relevant as system units can't contact the user dbus session anyways.

@LuanVSO
LuanVSO force-pushed the master branch 4 times, most recently from ef88696 to f8d52bd Compare September 14, 2026 16:25
@abraunegg

Copy link
Copy Markdown
Owner

@LuanVSO

Thanks for taking the earlier feedback into account and for dropping the change to onedrive@.service. I agree that the system-level templated service should not be ordered against the user/session D-Bus, so narrowing this PR to onedrive.service is a good change.

However, after reviewing the revised commit, I still cannot accept the current implementation.

The PR has now moved beyond replacing the startup delay with D-Bus ordering and changes the lifecycle of onedrive.service itself:

After=network-online.target dbus.socket graphical-session.target
Wants=network-online.target dbus.socket
PartOf=graphical-session.target

...

WantedBy=graphical-session.target

This creates a significant behavioural change.

graphical-session.target is specifically intended for user services which only apply while a graphical session exists. systemd documents PartOf=graphical-session.target as the mechanism for stopping such services when that graphical session terminates.

The OneDrive service is not a graphical-session-only service.

Desktop notifications are optional functionality, but the synchronisation service itself is explicitly supported in environments where there is no graphical session at all.

For example, the OneDrive documentation currently supports configuring a user service to start without user login using:

loginctl enable-linger <user>

That is specifically intended to allow the user's OneDrive service to start and continue operating independently of an interactive graphical login.

Changing:

WantedBy=default.target

to:

WantedBy=graphical-session.target

breaks that operating model for headless / non-graphical users.

Adding:

PartOf=graphical-session.target

also means the OneDrive service can be stopped when the graphical session terminates.

Your commit message explicitly describes this as:

this makes systemd stop the unit when the user logs out

That is not behaviour we want for the OneDrive synchronisation service.

Logging out of KDE/GNOME must not imply that a user's background OneDrive synchronisation service should stop, particularly when systemd lingering has intentionally been configured.

There is also an important distinction regarding:

After=graphical-session.target

graphical-session.target does not guarantee that the complete desktop environment has "fully initialized" or that all session services, notification services and other login activity have settled.

It is primarily a lifecycle/synchronisation target indicating that a graphical session exists.

Similarly, I agree that:

Wants=dbus.socket
After=dbus.socket

is a stronger relationship than After=dbus.socket alone, but this only addresses activation/ordering of that D-Bus socket. It still does not reproduce the broader startup grace period currently provided by the 15-second delay.

The existing delay was deliberately useful for more than D-Bus alone. It provides a short period for the wider boot/login environment to settle before OneDrive begins monitor and synchronisation processing.

I think we should solve the actual problem more narrowly

The systemd-analyze --user plot you supplied is useful and I accept the underlying issue: using:

ExecStartPre=/bin/sh -c 'sleep 15'

means the OneDrive service remains in its startup phase for those 15 seconds and can therefore delay completion of the user default.target.

However, I do not think solving that requires changing the service from a normal user background service into a graphical-session-scoped service.

The current service does not specify Type=, which means systemd uses Type=simple.

With Type=simple, systemd considers the service started once its main process has been created and then continues processing follow-up jobs.

I therefore think we should investigate a much more focused change:

ExecStart=/bin/sh -c 'sleep 15; exec @prefix@/bin/onedrive --monitor'

rather than:

ExecStartPre=/bin/sh -c 'sleep 15'
ExecStart=@prefix@/bin/onedrive --monitor

The shell becomes the initial main process, so systemd can consider the Type=simple service started immediately and allow default.target to continue.

After 15 seconds the shell uses exec to replace itself with the OneDrive process, retaining the same PID.

That potentially preserves all of the existing behaviour we actually require:

  • the intentional 15-second startup grace period remains;
  • default.target should no longer be delayed by an ExecStartPre;
  • WantedBy=default.target remains unchanged;
  • headless systems remain supported;
  • loginctl enable-linger behaviour remains supported;
  • OneDrive continues running independently of graphical login/logout;
  • no assumptions about graphical-session.target are introduced;
  • no assumption that D-Bus is the only subsystem which benefits from the startup delay is introduced.

This is the direction I would prefer to test before making further changes to the service dependency/lifecycle model.

In particular I would test the revised service with:

systemd-analyze --user plot
systemd-analyze --user critical-chain default.target

and confirm that:

  1. default.target is no longer delayed for 15 seconds;
  2. the OneDrive executable itself does not begin until approximately 15 seconds after activation;
  3. normal GUI notification behaviour remains available following graphical login;
  4. a user configured with loginctl enable-linger still starts OneDrive without an interactive login;
  5. logging out of a graphical session does not stop a deliberately lingering OneDrive service.

I think this would address the issue demonstrated by your systemd plot without changing the supported operating model of the OneDrive service.

@LuanVSO

LuanVSO commented Sep 14, 2026

Copy link
Copy Markdown
Author

@abraunegg what do you think of adding a systemd timer instead of using the shell trick?
in my experience changing the execstart messes with the service is identified in logs, using the timer would resolve all the concerns and also be applicable to the system version of the service

@abraunegg

abraunegg commented Sep 14, 2026 •

Copy link
Copy Markdown
Owner

@LuanVSO

Thanks for suggesting the timer approach.

I have considered it, and while I agree that a timer could technically avoid the current 15-second delay contributing to completion of default.target, I do not think introducing a timer is the right architectural direction for this project.

A timer would introduce a second activation mechanism purely to control when an existing long-running service starts.

That creates several additional support and lifecycle considerations which do not exist today.

For example:

  • existing installations currently enable onedrive.service directly;
  • changing to timer-based activation introduces an upgrade/migration problem because an already-enabled onedrive.service can continue to start directly and bypass the timer;
  • users with multiple OneDrive configurations commonly create additional/copied service units, which would then potentially require corresponding timer units;
  • troubleshooting now involves both a timer and a service rather than just the service;
  • systemctl start onedrive.service would bypass the timer entirely;
  • automatic service restarts would also bypass the timer unless additional machinery was introduced;
  • timer accuracy and expiry semantics become another configuration detail that needs to be understood and supported;
  • documentation for enabling, disabling, restarting and diagnosing the service would need to account for the additional unit.

That is a considerable amount of additional operational complexity to solve what is fundamentally a service-start ordering problem.

I would prefer that we keep the existing service model and solve the underlying problem directly.

The requirement that should be worked towards

The goal should be:

Start OneDrive at the latest portable point in normal user-manager startup, without tying it to a graphical login, and then retain only a short settling delay for things systemd cannot express as dependencies.

There are two important parts to that.

Firstly, OneDrive should start as late as we can reasonably and portably arrange within normal systemd --user startup.

The current service is pulled into:

WantedBy=default.target

and the existing 15-second ExecStartPre therefore causes OneDrive's startup job to remain outstanding while default.target is being reached.

That explains the behaviour shown in your systemd-analyze --user plot.

Rather than adding another activation mechanism, I think we should investigate whether the service can be ordered at a later appropriate synchronization point in the normal user-manager startup graph.

The objective is effectively:

user systemd startup
        |
        +-- sockets / D-Bus
        |
        +-- basic user environment
        |
        +-- normal user services
        |
        +-- normal user startup target reached
        |
        +-- small settling period
        |
        +-- OneDrive starts

rather than:

user systemd startup
        |
        +-- OneDrive becomes part of startup
              |
              +-- waits 15 seconds
              |
        +-- normal startup target finally completes

This must remain independent of graphical login

I specifically do not want to use graphical-session.target as that synchronization point.

OneDrive is not a graphical-session application.

The user service is also used on:

  • headless systems;
  • systems accessed only by SSH;
  • machines where GUI notifications are disabled;
  • systems configured with loginctl enable-linger;
  • systems where OneDrive is expected to continue synchronising after graphical logout.

So whatever point we select must remain valid for a normal user manager regardless of whether KDE, GNOME or any other graphical desktop session exists.

The graphical session may be one of the things that has finished starting by the time OneDrive begins on a normal desktop machine, but it must not become a prerequisite for the synchronization service itself.

The delay can then become much smaller

I also do not believe that the existing 15-second delay necessarily needs to remain 15 seconds if we move OneDrive later in the startup sequence.

The original delay provides a broad grace period for a number of things to settle:

  • D-Bus;
  • network availability;
  • user runtime environment;
  • desktop/login services where applicable;
  • other background services;
  • boot/login CPU and filesystem I/O.

If systemd ordering can already move OneDrive much closer to the end of normal user-manager startup, then most of that work should already have happened before OneDrive becomes eligible to run.

At that point the remaining delay is no longer responsible for allowing the whole user environment to initialise.

It becomes only a small final settling period for things which cannot be represented reliably through systemd dependencies.

For example, we may find through testing that:

3 seconds

or:

5 seconds

is sufficient once OneDrive is being started at the correct point.

I would rather establish that empirically across several distributions than preserve an arbitrary 15 seconds forever.

Proposed direction

I therefore suggest you pause the timer approach and first investigate the systemd user dependency graph properly.

The questions I think that need to be answered are:

  1. What is the latest portable synchronization point in normal systemd --user startup that exists for both graphical and headless/lingering users?
  2. Can onedrive.service be ordered after that point without creating a dependency cycle or changing normal shutdown behaviour?
  3. Can we preserve the current simple enable/disable model for onedrive.service?
  4. Once OneDrive is moved later, what is the smallest settling delay that remains reliable across Fedora, Debian, Ubuntu and other supported systemd distributions?
  5. Does this preserve loginctl enable-linger, headless operation, manual service control and multiple OneDrive service configurations?

I think that produces a substantially cleaner end result:

one service
one enablement model
one documented lifecycle
correct systemd ordering
small measured grace period

rather than adding a timer solely to work around where the service currently sits in the startup transaction.

So my preference is:

Do not add another systemd timer at this stage. First determine the latest portable point at which the existing OneDrive user service can start, then retain only the minimum settling delay that testing demonstrates is necessary.

I think that solves the actual problem while keeping the service architecture simple for users and maintainers.

@LuanVSO

LuanVSO commented Sep 14, 2026

Copy link
Copy Markdown
Author

nevermind, i didn't notice the exec on your previous suggestion, that will work

@LuanVSO LuanVSO changed the title systemd: wait on dbus.socket instead of sleeping systemd: prevent hanging default target Sep 14, 2026
@abraunegg

Copy link
Copy Markdown
Owner

@LuanVSO

Thanks for continuing to work through this.

I have re-reviewed the latest revision of this PR.

The current approach is now materially different from the original D-Bus-based proposal and from the later graphical-session.target approach.

The PR now effectively changes:

ExecStartPre=/bin/sh -c 'sleep 15'
ExecStart=@prefix@/bin/onedrive --monitor

to:

ExecStart=/bin/sh -c 'sleep 15; exec @prefix@/bin/onedrive --monitor'

and applies the same general approach to onedrive@.service.

This is a substantially cleaner and safer direction than introducing additional D-Bus dependencies, tying OneDrive to graphical-session.target, or adding a separate systemd timer.

There are, however, several points that need to be addressed before this can be considered for merge.

1. There is currently a functional regression in onedrive@.service

This is the immediate blocker.

The existing system service contains:

ExecStart=@prefix@/bin/onedrive --monitor --confdir=/home/%i/.config/onedrive

The current PR changes that to:

ExecStart=/bin/sh -c 'sleep 15; exec @prefix@/bin/onedrive --monitor'

and in doing so drops:

--confdir=/home/%i/.config/onedrive

That changes the behaviour of the templated system service and must be corrected.

At minimum, if this approach is retained, it should preserve the existing invocation:

ExecStart=/bin/sh -c 'sleep 15; exec @prefix@/bin/onedrive --monitor --confdir=/home/%i/.config/onedrive'

The --confdir option here is not incidental; it is part of how the onedrive@<user>.service template ensures that the correct user's OneDrive configuration is used.

2. The new mechanism does solve the specific default.target blocking problem

The current onedrive.service does not explicitly specify Type=, so systemd uses the normal Type=simple behaviour.

With:

ExecStart=/bin/sh -c 'sleep 15; exec @prefix@/bin/onedrive --monitor'

systemd starts /bin/sh as the service's main process and can consider the service started immediately.

The shell then waits for 15 seconds and finally uses exec to replace itself with the OneDrive process using the same PID.

Conceptually:

systemd starts onedrive.service
        |
        +-- /bin/sh becomes MainPID
        |
        +-- service considered started
        |
        +-- default.target can continue
        |
        +-- sleep 15
        |
        +-- exec onedrive --monitor
                |
                +-- same MainPID

So, for the specific problem demonstrated by your systemd-analyze --user plot, this should avoid having ExecStartPre hold the service in its activating state for 15 seconds and therefore avoid extending the time taken for default.target to be reached.

That part of the approach is technically reasonable.

3. However, this does not move OneDrive later in the systemd startup sequence

This distinction is important.

The requirement I previously described was:

Start OneDrive at the latest portable point in normal user-manager startup, without tying it to a graphical login, and then retain only a short settling delay for things systemd cannot express as dependencies.

The current revision does not yet implement that.

The service still uses:

After=network-online.target
Wants=network-online.target

and:

WantedBy=default.target

The dependency/start position of OneDrive is therefore fundamentally unchanged.

What has changed is how systemd perceives the service during the 15-second settling period.

Previously:

OneDrive service begins activation
        |
        +-- ExecStartPre sleep 15
        |
        +-- OneDrive starts
        |
        +-- service becomes started
        |
        +-- default.target can complete

The new behaviour is:

OneDrive service begins activation
        |
        +-- /bin/sh starts
        |
        +-- service is considered started
        |
        +-- default.target can complete
        |
        +-- shell sleeps 15 seconds
        |
        +-- shell execs OneDrive

That is an important improvement for the reported Plasma/systemd startup issue, but it is not the same as actually ordering OneDrive later in user-manager startup.

4. This changes what active (running) means during the first 15 seconds

This is probably the main supportability question with this approach.

For approximately 15 seconds after systemd starts the unit:

systemctl --user status onedrive.service

may legitimately show the service as:

active (running)

while the MainPID is still:

/bin/sh

and the OneDrive application itself has not yet started.

Similarly:

systemctl --user is-active onedrive.service

can return:

active

before onedrive --monitor has actually executed.

After the delay, exec replaces the shell with OneDrive using the same PID, which is good and avoids leaving a wrapper process around permanently.

However, we need to decide whether we are comfortable with this service-state semantic.

From a support perspective, somebody could run:

systemctl --user status onedrive.service

immediately after login and see:

Active: active (running)
Main PID: <pid> (sh)

even though the OneDrive client itself is still deliberately waiting to start.

That is not necessarily a reason to reject the approach, but it needs to be understood and tested because it changes what "service is active" represents during startup.

5. I do not think the 15-second delay should be reduced as part of this revision

We previously discussed whether the settling period could eventually be reduced from 15 seconds to something such as 3–5 seconds.

I still think that may be possible, but only if OneDrive is first moved later in the startup sequence.

This revision has not done that.

The application becomes eligible at essentially the same point as before; the only difference is that the settling period has moved from ExecStartPre into the main ExecStart command.

Therefore none of the assumptions behind the existing delay have materially changed.

The 15 seconds still allows time for:

  • user/system startup activity to settle;
  • D-Bus/session infrastructure to become usable;
  • network connectivity to become genuinely useful;
  • desktop/login services to initialise where applicable;
  • user runtime environment state to settle;
  • boot/login filesystem and CPU contention to reduce.

Until we have demonstrated a later and more deterministic start point, I would not shorten that safety margin.

6. I would still like the "latest portable startup point" question answered

The shell approach may ultimately prove to be the pragmatic solution.

There may simply not be a clean, portable systemd user target that means:

all normal user-manager startup has completed, but no graphical session is required, and this service can now start.

If achieving that requires:

  • DefaultDependencies=no;
  • dependency cycles around default.target;
  • graphical-session coupling;
  • additional timer units;
  • distribution-specific targets;
  • or other fragile systemd behaviour;

then I would prefer this simple shell approach over introducing that complexity.

However, I would like that conclusion to be reached deliberately.

In other words, before we settle on this approach, I would like confirmation that you have investigated whether there is a later portable user-manager synchronization point and concluded that doing so would either be unsafe, non-portable or materially more complex than the current solution.

If that is the case, then:

ExecStart=/bin/sh -c 'sleep 15; exec ...'

becomes a reasonable pragmatic answer to the specific problem.

7. Applying the same change to onedrive@.service also needs explicit validation

The original reported problem was the user service delaying user default.target / Plasma startup.

onedrive@.service is different:

WantedBy=multi-user.target

and runs as a system service for a specified user.

There is some logic in applying the same technique there because the existing ExecStartPre can similarly keep that service in its activation phase while multi-user.target is being reached.

However, this should not simply be assumed to behave correctly because the user service does.

Please explicitly test:

systemctl start onedrive@<user>.service

and verify that:

  • the correct user is used;
  • the correct /home/<user>/.config/onedrive configuration directory is used;
  • the service remains active after the shell performs the exec;
  • normal synchronisation operates correctly;
  • restart behaviour remains correct;
  • exit code 78 continues to honour RestartPreventExitStatus=78.

The fact that --confdir has already been accidentally lost in the current change reinforces why this path needs its own validation.

8. Restart behaviour should also be verified

One useful property of this approach is that automatic restarts should continue to receive the same settling period.

Today the sequence is effectively:

OneDrive fails
        |
        +-- RestartSec=3
        |
        +-- ExecStartPre sleep 15
        |
        +-- OneDrive starts

With the proposed change:

OneDrive fails
        |
        +-- RestartSec=3
        |
        +-- /bin/sh starts
        |
        +-- sleep 15
        |
        +-- exec OneDrive

So that behaviour is broadly preserved.

Because exec replaces the shell rather than launching OneDrive as a child process, the same MainPID is retained and the eventual exit status remains associated with the service's main process.

That should mean the existing:

RestartPreventExitStatus=78

behaviour remains intact, but I would still like this explicitly tested.

9. Testing required for this revision

If this implementation is retained, I would like to see the following validation.

For the user service:

systemctl --user daemon-reload
systemctl --user restart onedrive.service
systemctl --user status onedrive.service

During the first 15 seconds, capture:

systemctl --user show onedrive.service \
    -p ActiveState \
    -p SubState \
    -p MainPID \
    -p ExecMainPID

and inspect:

ps -fp <MainPID>

Then repeat after the 15-second delay to demonstrate the transition from:

/bin/sh

to:

onedrive

using the same PID.

Please also test:

  • actual boot/login startup;
  • Fedora;
  • Debian;
  • preferably at least one other mainstream systemd distribution;
  • graphical login;
  • loginctl enable-linger with no graphical login;
  • logout while lingering is enabled;
  • systemctl --user start;
  • systemctl --user stop;
  • systemctl --user restart;
  • deliberate OneDrive failure followed by automatic restart;
  • exit code 78 / resync-required behaviour;
  • GUI notification availability once OneDrive actually starts.

For the system service, please also test:

systemctl start onedrive@<user>.service
systemctl status onedrive@<user>.service

and verify the correct:

/home/<user>/.config/onedrive

configuration is being used.

The most useful systemd comparison would also be before/after output from:

systemd-analyze --user critical-chain default.target
systemd-analyze --user plot

to demonstrate that the original 15-second default.target delay has actually been eliminated.

10. The PR description now also needs updating

The PR title has been updated to:

systemd: prevent hanging default target

which better reflects the current objective.

However, the PR description still refers to:

After=dbus.socket

which is no longer part of the implementation.

Once the technical approach is finalised, please update the PR description so it describes the actual problem and solution being proposed.


So where I currently land is:

The latest revision is substantially better than the earlier D-Bus, graphical-session and timer proposals.

I think the shell + exec mechanism is technically capable of solving the specific problem demonstrated by the systemd-analyze output while retaining the existing 15-second settling period and without introducing another systemd unit.

However, before this is mergeable:

  1. the lost --confdir=/home/%i/.config/onedrive in onedrive@.service must be restored;
  2. the changed active (running) semantics during the 15-second startup window need to be explicitly accepted and tested;
  3. both the user and system service paths need validation;
  4. the investigation into whether there is a later portable user-manager synchronization point should be concluded;
  5. the existing 15-second period should remain unchanged unless later startup ordering is actually introduced and independently validated;
  6. the PR description should be updated to match the implementation.

If the systemd dependency investigation demonstrates that there is no cleaner portable "late startup" point without introducing substantially greater complexity, then I would be open to treating this shell + exec approach as the pragmatic solution.

Adjust service startup ordering so OneDrive does not delay session
initialization while waiting for the default or multi-user target.
@LuanVSO

LuanVSO commented Sep 15, 2026

Copy link
Copy Markdown
Author

to answer your question, all the distros that i tested(fedora,arch,debian,ubuntu) activate dbus on basic.target that is always reached before default.target

after thinking over this some more, i think I've reached a cleaner solution that keeps your delay, doesn't delay session init and doesn't change any semantics about the service.

Update PR with solution, tested with systemd 257
@abraunegg

Copy link
Copy Markdown
Owner

@LuanVSO

Please can you test the proposed change I have just made.

@LuanVSO

LuanVSO commented Sep 16, 2026

Copy link
Copy Markdown
Author

can confirm it fixes the issue thanks, the desktop now runs all of its services early, and shutting the machine down does not take longer than expected.

new systemd-analize plot user

@abraunegg abraunegg added this to the v2.5.12 milestone Sep 17, 2026
@abraunegg
abraunegg marked this pull request as ready for review September 17, 2026 04:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants