Dalarna University's logo and link to the university's website

du.sePublications
Change search
Link to record
Permanent link

Direct link
Hintze, Arend, ProfessorORCID iD iconorcid.org/0000-0002-4872-1961
Publications (10 of 87) Show all publications
Halabi, R., Mulsant, B. H., Tolend, M., Blumberger, D. M., DeShaw, A., Hintze, A., . . . Ortiz, A. (2026). A systematic exploration of digital biomarkers for the detection of depressive episodes in bipolar disorder. Npj mental health research, 5(1), Article ID 13.
Open this publication in new window or tab >>A systematic exploration of digital biomarkers for the detection of depressive episodes in bipolar disorder
Show others...
2026 (English)In: Npj mental health research, ISSN 2731-4251, Vol. 5, no 1, article id 13Article in journal (Refereed) Published
Abstract [en]

Digital phenotyping promises to transform psychiatry by using multimodal, densely sampled data. However, its potential is hindered by the lack of focus on identifying and validating digital biomarkers that accurately reflect mental states before evaluating their impact on outcomes. This longitudinal study used explainable machine learning to analyze multivariate, densely sampled data from 133 bipolar disorder (BD) participants over a median of 251 days, identifying robust digital biomarkers defining depressive episodes. The analysis included features from email-based daily self-reported mood, energy, and anxiety, as well as passively collected activity and sleep data using an Oura ring. The most robust descriptors of depressive episodes were lower daily mood variability, lower daily activity variability, and higher daily sleep onset latency variability. Self-reported daily mood features achieved the highest performance (AU-ROC: 0.82 ± 0.03). Our results establish the value of multimodal data and represent a critical first step toward automated detection and prediction of illness episodes in BD.

National Category
Psychiatry
Identifiers
urn:nbn:se:du-53083 (URN)10.1038/s44184-026-00195-5 (DOI)001696228000001 ()41724810 (PubMedID)2-s2.0-105030608668 (Scopus ID)
Available from: 2026-02-24 Created: 2026-02-24 Last updated: 2026-03-11Bibliographically approved
Hintze, A., Proschinger Åström, F. & Schossau, J. (2026). Autonomous language-image generation loops converge to generic visual motifs. Patterns, 7(1), Article ID 101451.
Open this publication in new window or tab >>Autonomous language-image generation loops converge to generic visual motifs
2026 (English)In: Patterns, E-ISSN 2666-3899, Vol. 7, no 1, article id 101451Article in journal (Refereed) Published
Abstract [en]

Autonomous AI-to-AI creative systems promise new frontiers in machine creativity, yet we show that they systematically converge toward generic outputs. We built iterative feedback loops between Stable Diffusion XL (SDXL; image generation) and Large Language and Vision Assistant (LLaVA; image description), forming autonomous text → image → text → image cycles. Across 700 trajectories with diverse prompts and 7 temperature settings over 100 iterations, all runs converged to nearly identical visuals—what we term ‘‘visual elevator music.’’ Quantitative analysis revealed just 12 dominant motifs with commercially safe aesthetics, such as stormy lighthouses and palatial interiors. This convergence persisted across model pairs, indicating structural limits in cross-modal AI creativity. The effect mirrors human cultural transmission, where iterated learning amplifies cognitive biases, but here, diversity collapses entirely as AI loops gravitate to high-probability attractors in training data. Our findings expose hidden homogenizing tendencies in current architectures and underscore the need for anti-convergence mechanisms and sustained human-AI interplay to preserve creative diversity.

Place, publisher, year, edition, pages
Cell Press, 2026
Keywords
Computational linguistics; Dynamical systems; Iterative methods; Learning systems; AI systems; Attractor dynamics; Autonomous AI system; Computational creativities; Convergence; Drift; Generative AI; Language model; Large language model; Vision-language model; Visual languages
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:du-52292 (URN)10.1016/j.patter.2025.101451 (DOI)001664473200002 ()41583980 (PubMedID)2-s2.0-105025359872 (Scopus ID)
Available from: 2026-01-16 Created: 2026-01-16 Last updated: 2026-02-19Bibliographically approved
Proschinger Åström, F. & Hintze, A. (2026). How to systematically and quantifiably remove meaning?. Frontiers in Artificial Intelligence, 9, Article ID 1783410.
Open this publication in new window or tab >>How to systematically and quantifiably remove meaning?
2026 (English)In: Frontiers in Artificial Intelligence, E-ISSN 2624-8212, Vol. 9, article id 1783410Article in journal (Refereed) Published
Abstract [en]

Large language models increasingly mediate real-world tasks, yet we lack systematic ways to quantify how their performance degrades when the meaning of their inputs is eroded. To bridge this gap, we developed a framework to semantically erode meaning and quantify its intensity, grounded in discourse analysis, psycholinguistics, and software engineering, comprising five theoretically motivated methods: omission of key information and context, lexical substitution with near-synonyms, increased abstraction, structural obfuscation and renaming, and injection of logical errors. We applied these erosion operators across five domains and quantified their effects on model performance using a publicly available language model. A two-way Analysis of Variance (ANOVA) revealed significant main effects of both domain and erosion method, as well as a significant interaction, indicating that the impact of semantic degradation depends jointly on how text is eroded and how domain-specific information is encoded. Logical error erosions proved especially damaging for code generation, whereas structural obfuscation most strongly impaired news and instruction tasks. Epistasis analysis of pairwise erosion unions showed that some combinations produced super-additive degradation while others exhibited compensatory effects. These domain-by-erosion profiles provide diagnostic insight into where multi-step large language model (LLM) pipelines are most likely to fail and suggest that robustness benchmarks should probe models along domain-specific vulnerability dimensions rather than relying on generic perturbations. Semantic erosion thus offers a principled tool for turning model failure into evidence about how language models structure and degrade meaning.

Keywords
large language models; meaning and semantics; meaning degradation; robustness evaluation; semantic erosion.
National Category
Natural Language Processing Software Engineering
Identifiers
urn:nbn:se:du-53781 (URN)10.3389/frai.2026.1783410 (DOI)001776866800001 ()42211171 (PubMedID)2-s2.0-105041462539 (Scopus ID)
Available from: 2026-06-02 Created: 2026-06-02 Last updated: 2026-06-23Bibliographically approved
Tolend, M., Halabi, R., Ghaouari, K., Lau, Y. C., Alda, M., Hintze, A., . . . Ortiz, A. (2026). Novel Abstract Screening Algorithm Using Delphi-Inspired Large Language Model Consensus for Systematic Reviews in Psychiatry: Nouvel algorithme de sélection des résumés utilisant un consensus issu d’un grand modèle de langage inspiré de la méthode Delphi pour les revues systématiques en psychiatrie. Canadian journal of psychiatry
Open this publication in new window or tab >>Novel Abstract Screening Algorithm Using Delphi-Inspired Large Language Model Consensus for Systematic Reviews in Psychiatry: Nouvel algorithme de sélection des résumés utilisant un consensus issu d’un grand modèle de langage inspiré de la méthode Delphi pour les revues systématiques en psychiatrie
Show others...
2026 (English)In: Canadian journal of psychiatry, ISSN 0706-7437Article in journal (Refereed) Epub ahead of print
Abstract [en]

Background: Large language models (LLMs) may reduce the burden associated with performing systematic reviews by prescreening abstracts from a literature search for eligibility for inclusion in full-text review. Methods: We developed an iterative, LLM-based workflow for screening abstracts: after manual specification of eligibility criteria and seed examples, an ensemble of five LLMs deliberates through a Delphi process to classify a batch of abstracts; these labels are used to train a logistic regression model that ranks the remaining abstracts and identifies a new batch of abstracts for LLM escalation until all abstracts are labelled by the LLM or probability thresholds. We tested our workflow on abstracts screened in three published systematic reviews in psychiatry. Our primary endpoint was the recall metric, and secondary endpoint was the work saved over sampling at 95% recall metric (WSS@95%). Results: In a dataset on autism biomarkers, 1,655 (35%) of 4,745 retrieved abstracts were judged to be relevant by the original authors. The Delphi–LLM workflow correctly identified 1,605 (97.0%) of these 1,655 abstracts (precision = 54.2%, WSS@95% = 38.1%). The performance metrics were better than non-LLM approaches (recall ≤ 91%, WSS@95 ≤ 26%), and, overall, balanced these metrics optimally compared to single-LLM agents (recall = 84.9–99.9%, WSS@95% = 16.7–39.8%). The recall and work saved metrics were similarly reliable and among the top in two low-prevalence datasets on an attention-deficit hyperactivity disorder treatment review (10% of 2,891 relevant) and a posttraumatic stress disorder trajectory review (7% of 4,453 relevant). For these two datasets, recall was 100.0% and 96.4%, and the WSS@95% was 17.3% and 18.5%, respectively. Conclusions: We presented the design and validation of a novel abstract screening workflow that centres around a Delphi-style aggregation process to harness the strengths of five open-source LLMs that can be run on consumer-level workstations. This multi-LLM workflow showed acceptable and reliable performance for use as an automated prescreening method to facilitate systematic reviews. © The Author(s) 2026. This article is distributed under the terms of the Creative Commons Attribution-NonCommercial 4.0 License (https://creativecommons.org/licenses/by-nc/4.0/) which permits non-commercial use, reproduction and distribution of the work without further permission provided the original work is attributed as specified on the SAGE and Open Access page (https://us.sagepub.com/en-us/nam/open-access-at-sage).

Place, publisher, year, edition, pages
SAGE Publications Inc., 2026
Keywords
abstract screening, Delphi method, large language models, systematic reviews, text embedding
National Category
Computer and Information Sciences Clinical Medicine
Identifiers
urn:nbn:se:du-53643 (URN)10.1177/07067437261445767 (DOI)001753735500001 ()42059529 (PubMedID)2-s2.0-105037549282 (Scopus ID)
Available from: 2026-05-11 Created: 2026-05-11 Last updated: 2026-05-19Bibliographically approved
Ortiz, A., Halabi, R., Blumberger, D., Gonzalez-Torres, C., Hintze, A., Husain, M. I., . . . Mulsant, B. H. (2026). Trajectories of Suicidal Risk Impact Mood Regulation Differently in Patients With a Diagnosis of Bipolar Disorder. Acta Psychiatrica Scandinavica, 154(1), 60-73
Open this publication in new window or tab >>Trajectories of Suicidal Risk Impact Mood Regulation Differently in Patients With a Diagnosis of Bipolar Disorder
Show others...
2026 (English)In: Acta Psychiatrica Scandinavica, ISSN 0001-690X, E-ISSN 1600-0447, Vol. 154, no 1, p. 60-73Article in journal (Refereed) Published
Abstract [en]

BACKGROUND: Bipolar disorder (BD) carries a suicide risk 20 times higher than the general population, with up to 60% of patients attempting suicide. Current interventions have failed to reduce its incidence; static factors have shown limited predictive utility. Emerging evidence suggests dynamic monitoring approaches may offer complementary value. This study examined whether quantifiable differences in mood regulation patterns exist across the suicidality continuum among patients diagnosed with BD.

METHOD: We analyzed daily self-reported mood, anxiety, and energy levels from 164 participants recruited from two Canadian academic hospitals (April 2021-August 2024). Participants were stratified into six groups based on suicide attempt history, current polarity, and active suicidality status. Using time-series analysis, we computed autocorrelation and cross-correlation functions to examine temporal relationships within and between variables across 1-7 day lags. Data comprised 64,351 valid observations over 461.5 ± 236.6 days of follow-up.

RESULTS: Participants with the highest suicide risk (previous attempt, in a current depressive episode with active suicidality) demonstrated significantly higher day-to-day autocorrelation compared to the lowest-risk participants (no prior attempts and currently euthymic) for mood (0.53 vs. 0.29, p = 0.01), energy (0.52 vs. 0.23, p = 0.02), and anxiety series (0.55 vs. 0.32, p = 0.04). Cross-correlation analysis revealed mood-energy decoupling during active suicidality; as well as a stronger negative mood-anxiety correlation in those with a prior attempt, even during euthymia.

CONCLUSION: Higher autocorrelation patterns are indicative of a pathologically stable mood regulation in high-risk individuals, potentially serving as dynamic biomarkers for suicide risk stratification and targeted intervention development. Our findings demonstrate that a more aggressive approach to treating comorbid anxiety may be essential for reducing the risk of future attempts. They also challenge traditional conceptualizations equating euthymia with the absence of suicide risk, suggesting neurobiological vulnerability despite symptomatic remission.

Keywords
bipolar disorder, mood regulation, suicidal behavior, time‐series analysis
National Category
Psychiatry
Identifiers
urn:nbn:se:du-53294 (URN)10.1111/acps.70094 (DOI)001731014400001 ()41924933 (PubMedID)2-s2.0-105034861601 (Scopus ID)
Available from: 2026-04-09 Created: 2026-04-09 Last updated: 2026-06-16Bibliographically approved
Bohm, C., Adami, C. & Hintze, A. (2026). What Is Redundancy?. Entropy, 28(2), Article ID 167.
Open this publication in new window or tab >>What Is Redundancy?
2026 (English)In: Entropy, E-ISSN 1099-4300, Vol. 28, no 2, article id 167Article in journal (Refereed) Published
Abstract [en]

Redundancy is a central yet persistently ambiguous concept in multivariate information theory. Across the literature, the same term is used to describe fundamentally distinct phenomena. Operational redundancy concerns how different inputs relate to the prediction of output states, while informational redundancy concerns content overlap among inputs relevant to an output. These notions are routinely conflated in decompositions of mutual information, leading to incompatible definitions, contradictory interpretations, and apparent paradoxes-particularly when inputs are statistically independent. We argue that the difficulty in defining redundancy is not primarily technical, but conceptual: the field has not converged on what redundancy is meant to signify. We formalize this distinction by identifying two classes of redundancy. Operational redundancy encompasses task-relative properties and covers conditions when inputs are sufficient or substitutable for prediction. Informational redundancy concerns shared content among inputs, grounded in mutual information between them. Using functional examples and biased input ensembles, we demonstrate the practical distinction between these classes: inputs with no informational overlap can exhibit operational redundancy, while partial observation can induce statistical correlations that create content overlap without reflecting the underlying functional structure. We conclude by proposing a clear separation of these concepts and outlining minimal commitments for each. This separation clarifies why redundancy remains elusive, why no single measure can satisfy all intuitions, and how future work can proceed without redefining information itself.

Keywords
information decomposition, redundancy, synergy
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:du-53119 (URN)10.3390/e28020167 (DOI)41751670 (PubMedID)2-s2.0-105031280308 (Scopus ID)
Available from: 2026-03-03 Created: 2026-03-03 Last updated: 2026-03-09Bibliographically approved
Ortiz, A., Halabi, R., Tolend, M., Gonzalez-Torres, C., Blumberger, D. M., Husain, I., . . . Mulsant, B. (2025). A Systematic Exploration of Which Digital Biomarkers are the Most Accurate to Detect Depressive Episodes in Bipolar Disorder. Bipolar Disorders, 27(supp. 1), S147-S147
Open this publication in new window or tab >>A Systematic Exploration of Which Digital Biomarkers are the Most Accurate to Detect Depressive Episodes in Bipolar Disorder
Show others...
2025 (English)In: Bipolar Disorders, ISSN 1398-5647, E-ISSN 1399-5618, Vol. 27, no supp. 1, p. S147-S147Article in journal, Meeting abstract (Refereed) Published
Place, publisher, year, edition, pages
WILEY, 2025
National Category
Psychiatry
Identifiers
urn:nbn:se:du-51641 (URN)001578710100128 ()
Available from: 2025-11-03 Created: 2025-11-03 Last updated: 2025-11-03Bibliographically approved
Halabi, R., Tolend, M., Alda, M., Blumberger, D. M., Husain, I., O’Donovan, C., . . . Ortiz, A. (2025). Age-and Polarity-Related Changes in REM and NREM Sleep Architectures in Bipolar Disorder Using Densely Sampled Wearable Data. Paper presented at The 27th Annual Conference of the International Society for Bipolar Disorders, Chiba, Japan, 17 – 19 September 2025. Bipolar Disorders, 27(S1), 53-53
Open this publication in new window or tab >>Age-and Polarity-Related Changes in REM and NREM Sleep Architectures in Bipolar Disorder Using Densely Sampled Wearable Data
Show others...
2025 (English)In: Bipolar Disorders, ISSN 1398-5647, E-ISSN 1399-5618, Vol. 27, no S1, p. 53-53Article in journal, Meeting abstract (Refereed) Published
Abstract [en]

Introduction: This study aimed to systematically characterize differences in sleep architecture across clinical states and age groups in bipolar disorder (BD).

Method: We recruited 178 adults diagnosed with BD I or II, followed for a mean of 482 ± 243 days. Participants wore an Oura Ring to continuously collect a 5-min sampled hypnogram. Weekly Patient Health Questionnaire (PHQ-9) and Altman Self-Rating Mania Scale (ASRM) scores defined clinical states [euthymia, depression, (hypo)mania]. Sleep variables (e.g., light sleep duration) were compared across clinical states and age groups (median split: 37 years) using Kruskal–Wallis and pairwise Mann–Whitney U-tests with bootstrapping.

Results: We collected 5,834,600 hypnogram observations: 5,550,860 during euthymia, 261,504 during depressive episodes, and 22,236 during (hypo)manic episodes. (Hypo)mania was characterized by irregular wakefulness (i.e., fragmented sleep), the shortest REM sleep (1.41 ± 0.80 h), deep sleep (0.94 ± 0.75 h), and sleep onset latency (0.21 ± 0.24 h). Depression showed the longest deep sleep (1.82 ± 0.85 h), REM sleep (1.61 ± 0.93 h), and sleep onset latency (0.26 ± 0.31 h). During (hypo)mania, older adults (> 37 years) had shorter deep sleep (0.94 ± 0.75 h) and longer light sleep (4.67 ± 1.40 h) than younger adults.

Conclusion: This study provides a deeper analysis of sleep architecture in BD, leveraging objective, continuous data collected longitudinally using wearables. Our findings show that each clinical polarity is characterized by a distinct sleep architecture, which also differs across age groups. Further investigation into sleep phases may offer insights for treatment strategies. Age-related differences in sleep architecture were also observed, which do not mirror the patterns seen in healthy individuals.

Place, publisher, year, edition, pages
WILEY, 2025
National Category
Psychiatry
Identifiers
urn:nbn:se:du-51640 (URN)10.1111/bdi.70044 (DOI)001578710100229 ()
Conference
The 27th Annual Conference of the International Society for Bipolar Disorders, Chiba, Japan, 17 – 19 September 2025
Available from: 2025-10-31 Created: 2025-10-31 Last updated: 2025-10-31Bibliographically approved
Mehra, P. & Hintze, A. (2025). Continuous Evolution in the NK Treadmill Model. Artificial Life, 31(3), 256-275
Open this publication in new window or tab >>Continuous Evolution in the NK Treadmill Model
2025 (English)In: Artificial Life, ISSN 1064-5462, E-ISSN 1530-9185, Vol. 31, no 3, p. 256-275Article in journal (Refereed) Published
Abstract [en]

The NK fitness landscape is a well-known model with which to study evolutionary dynamics in landscapes of different ruggedness. However, the model is static, and genomes are typically small, allowing observations over only a short adaptive period. Here we introduce an extension to the model that allows the experimenter to set the velocity at which the landscape changes independently from other parameters, such as the ruggedness or the mutation rate. We find that, similar to the previously observed complexity catastrophe, where evolution comes to a halt when environments become too complex due to overly high degrees of epistasis, here the same phenomenon occurs when changes happen too rapidly. Our expanded model also preserves essential properties of the static NK landscape, allowing for proper comparisons between static and dynamic landscapes.

Keywords
Fitness landscape, dynamic landscape, epistasis, pleiotropy, ruggedness, velocity
National Category
Evolutionary Biology
Identifiers
urn:nbn:se:du-50247 (URN)10.1162/artl_a_00467 (DOI)001566839600001 ()39964771 (PubMedID)2-s2.0-105015685600 (Scopus ID)
Available from: 2025-02-26 Created: 2025-02-26 Last updated: 2025-10-31Bibliographically approved
Ortiz, A., Halabi, R., Alda, M., DeShaw, A., Husain, M. I., Nunes, A., . . . Hintze, A. (2025). Day-to-day variability in activity levels detects transitions to depressive symptoms in bipolar disorder earlier than changes in sleep and mood. International Journal of Bipolar Disorders, 13(1), Article ID 13.
Open this publication in new window or tab >>Day-to-day variability in activity levels detects transitions to depressive symptoms in bipolar disorder earlier than changes in sleep and mood
Show others...
2025 (English)In: International Journal of Bipolar Disorders, E-ISSN 2194-7511, Vol. 13, no 1, article id 13Article in journal (Refereed) Published
Abstract [en]

Anticipating clinical transitions in bipolar disorder (BD) is essential for the development of clinically actionable predictions. Our aim was to determine what is the earliest indicator of the onset of depressive symptoms in BD. We hypothesized that changes in activity would be the earliest indicator of future depressive symptoms. The study was a prospective, observational, contactless study. Participants were 127 outpatients with a primary diagnosis of BD, followed up for 12.6 (5.7) [(mean (SD)] months. They wore a smart ring continuously, which monitored their daily activity and sleep parameters. Participants were also asked to complete weekly self-ratings using the Patient Health Questionnaire (PHQ-9) and Altman Self-Rating Mania Scale (ASRS) scales. Primary outcome measures were depressive symptom onset detection metrics (i.e., accuracy, sensitivity, and specificity); and detection delay (in days), compared between self-rating scales and wearable data. Depressive symptoms were labeled as two or more consecutive weeks of total PHQ-9 > 10, and data-driven symptom onsets were detected using time-frequency spectral derivative spike detection (TF-SD2). Our results showed that day-to-day variability in the number of steps anticipated the onset of depressive symptoms 7.0 (9.0) (median (IQR)) days before they occurred, significantly earlier than the early prediction window provided by deep sleep duration (median (IQR), 4.0 (5.0) days; p <.05). Taken together, our results demonstrate that changes in activity were the earliest indicator of depressive symptoms in participants with BD. Transition to dynamic representations of behavioral phenomena in psychiatry may facilitate episode forecasting and individualized preventive interventions.

Keywords
Activity, Bipolar disorder, Densely-sampled, Mood variability, Onset, Sleep, Wearable technology
National Category
Psychiatry
Identifiers
urn:nbn:se:du-50444 (URN)10.1186/s40345-025-00379-6 (DOI)001458241700001 ()40175826 (PubMedID)2-s2.0-105001723805 (Scopus ID)
Available from: 2025-04-09 Created: 2025-04-09 Last updated: 2025-10-09Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-4872-1961

Search in DiVA

Show all publications