Abstract
Legislative knowledge evolves as an intricate hypertext in which documents are interconnected through complex, often implicit relationships. In this paper, we introduce ReSB2, a framework for retrieving and linking similar legislative bills that supports human–machine collaboration and helps reduce redundancy in the lawmaking process. The framework fine-tunes two ModernBERT-based language models on authentic legislative data, incorporating domain-specific formatting and procedural constraints derived from real workflows in a Brazilian state-level legislative assembly. To ensure transparency, ReSB2 integrates an explainability module based on Integrated Gradients, enabling analysts to inspect which textual elements most influence model decisions. Evaluated on a large corpus of official bills, the framework outperforms both general-purpose and domain-specific baselines in identifying semantically similar documents, achieving recall values of approximately 0.9. Human-centric evaluation with domain experts further demonstrates that ReSB2 serves as an effective human-centered augmentation tool, supporting the consistency and governance of legislative knowledge.
1 Introduction
Legislative processes involve the creation, review, and approval of bills that propose new laws or amendments to existing ones [6, 9]. These documents do not exist in isolation. Rather, they form a dynamic, non-linear network of information, interconnected through explicit citations and implicit semantic relationships. Figure 1 illustrates examples of bills and their hypertextual links based on similarity (linked) and keywords (indexing). Therefore, legal documents exhibit an inherent hypertextual nature [8, 10, 34].
In the Legislative Assembly of Minas Gerais (ALMG), a Brazilian state-level legislative assembly, the lawmaking workflow requires that new proposals be constantly integrated into this evolving network. When a new bill is introduced, procedural rules require identifying similarities to existing documents. If such a connection is found, a formal link is established, associating the documents to prevent redundancy and ensure the coherence of the legal ecosystem [6]. Importantly, similarity in this context is not defined solely by lexical overlap or textual resemblance. Instead, ALMG analysts establish similarity links when two bills address substantially the same legislative matter or pursue equivalent regulatory objectives, such that they should be analyzed together to avoid redundancy or inconsistent legal outcomes. This assessment relies on expert legal interpretation and is subsequently constrained by procedural rules defined in the ALMG Internal Rules.
Figure 1 illustrates this process, showing a bill linked by similarity to another, which is in turn connected to three additional bills, forming a non-linear network of information with five linked documents in total. However, navigating this intricate web of legal knowledge poses a significant scaling challenge. Identifying similar legislative bills is far from trivial: legislative language is highly technical, featuring complex sentence structures, frequent references to legal norms, and substantial variation in writing style.
This work addresses a domain-specific Semantic Textual Similarity (STS) task in which similarity is defined by legislative substance. Specifically, two bills are considered similar when they address substantially the same legislative matter or pursue equivalent regulatory objectives, even if they differ considerably in wording. Consequently, high-recall semantic retrieval must be combined with institutional constraints and human validation.
This challenge is further compounded by the fact that similarity identification is not merely a linguistic problem, but also a task of discovering associative trails governed by domain-specific rules and procedural constraints. For instance, the ALMG Internal Rules [6] impose temporal and situational filters for a similarity link to be considered valid, Article 173, Section III restricts comparisons to bills currently under discussion, whereas Article 284, Section I limits comparisons to bills that were approved or rejected within the same legislative session. Effectively transforming simple retrieval into a form of rule-aware navigation.
In this paper, we introduce ReSB2, a framework for machine-assisted linking of legislative texts that supports domain specialists in identifying similar bills in accordance with procedural requirements. The system acts as an augmentation tool, enabling effective human–machine collaboration by combining high-recall retrieval with interpretable similarity signals, allowing analysts to validate and refine suggested links.
ReSB2 utilizes domain-adapted ModernBERT models to identify semantic connections beyond the limitations of keyword-based tools. The framework integrates an Explainable Artificial Intelligence (XAI) module based on Integrated Gradients, allowing analysts to inspect the textual features that contribute to each suggested link. Validated through retrieval metrics and qualitative feedback from ALMG specialists, ReSB2 demonstrates how AI can serve as a stewardship tool for large-scale digital corpora, fostering consistency and transparency in the democratic process.
The main contributions of this paper are:
2 Related Work
The intersection of Legal Information Retrieval (LIR) and hypertext has a rich history, rooted in early legal retrieval systems [2, 25] and specialized legal hypertext databases [35]. More recently, the convergence of LIR and Artificial Intelligence has advanced the study of law, framing it as a complex and dynamic hypertextual ecosystem [10, 34].
The conceptualization of lawmaking as a hypertextual endeavor argues that legal texts are not merely static documents but dynamic structures that undermine traditional, linear theories of legislation and interpretation [8]. This paradigm shift has led to the development of integrated legal decision-support systems that prioritize human augmentation over autonomous robot lawyers, emphasizing the importance of sustainable, accessible legal advisory infrastructures [14]. Modern applications of this hypertextual approach extend to the virtualization of parliamentary debates, where systems like VR-ParlExplorer utilize Large Language Models (LLMs) and chatbots to create immersive, spatial hypertext environments for collaborative interaction [1]. Furthermore, the challenge of navigating these vast institutional corpora is addressed through AI-driven automatic indexing, which cross-references multimedia streams with verbatim reports to generate "Video Tables of Contents" (VTOCs), thereby enhancing the accessibility and non-linear navigability of public records [7]. Together, these works position hypertext as a suitable architecture for transparent, interlinked, and human-centered legal systems.
Operationalizing such hypertextual structures, however, requires robust semantic representations capable of capturing nuanced relationships between documents. This need has driven the adoption of Natural Language Processing (NLP) techniques, particularly in light of open government data policies that promote transparency and accessibility of public information [12, 18, 27].
For Portuguese-language applications, most models are derived from BERTimbau [29], a general-purpose model pre-trained on a large corpus of Portuguese texts collected from the web and based on BERT [11].
These models have been adapted to several specialized domains, including legislative and legal contexts. For example, GovBERT-BR [28] extends BERTimbau through continued pre-training on texts from the legal and administrative domains, improving performance on government-related NLP tasks. Another noteworthy contribution is JurisBERT [32], a BERT model pre-trained from scratch on Brazilian legal texts and refined for the legal sub-language. It was trained on pairs of legal case summaries annotated for Semantic Textual Similarity (STS). BERTikal [23] is a Brazilian legal-domain language model based on the BERT-base architecture, initialized from the BERTimbau [29] checkpoint and further pre-trained on corpora composed of Brazilian legal texts, this additional domain-specific training enables the model to better capture the linguistic patterns and terminology characteristic of legal documents.
More recently, ModernBERT [33] was introduced as an evolution of the BERT architecture, designed to handle long documents and extended contexts more efficiently. While maintaining the foundational design of BERT, ModernBERT incorporates key improvements such as an expanded context window and optimized attention mechanisms, allowing it to process large texts without substantial increases in computational cost. Building on this, Multilingual ModernBERT (mmBERT) [20] extended the architecture to support over 1,800 languages, including Portuguese. To the best of our knowledge, there has not yet been an adaptation of mmBERT to the Portuguese legal or legislative domain.
Alongside these models, several legislative datasets have been published. The UlyssesNER-Br dataset [3] contains documents from the Brazilian Chamber of Deputies annotated with named entities. LegisPL-BR [19] is another dataset that provides legislative texts from the same institution. Additionally, the dataset presented in [4] includes bills from the Rio Grande do Norte Parliament (ALRN), thereby extending coverage to a Brazilian state-level legislative assembly.
In Explainable Artificial Intelligence (XAI), recent work has applied Integrated Gradients (IG) [30] to improve the interpretability of BERT-based models. Talebi et al. [31] used IG to analyse decision-making in neuroradiology protocol assignment. Makino et al. [17] examined the sensitivity of IG to the number of integration steps and proposed adaptive step selection to reduce interpretability errors. These works show that IG-based explanations are reliable and semantically coherent, supporting their use in sensitive, high-stakes domains such as law and government.
3 SB2 Dataset: Similar Brazilian State Bills
Before introducing the proposed framework, this section describes the SB2 dataset and the preprocessing pipeline applied, both essential components of our approach. The dataset consists of legislative bill texts collected from two primary sources: the Legislative Assembly of Minas Gerais (ALMG) and the Brazilian Chamber of Deputies1.
Although other datasets include Brazilian legislative bills [3, 4, 19], none of them provide information on similarity relationships between bills. This is precisely the distinctive contribution of the SB2 dataset.
The preprocessing procedure applied to both sources involved several steps designed to reduce textual noise and remove content irrelevant to the similarity detection task:
Length-based filtering. Bills with fewer than 100 characters or with empty or null text were removed to eliminate incomplete or irrelevant records.
Removal of justification sections. Sections explicitly marked as “justification” (and their variants) were removed using regular expressions. Since these typically contain subjective arguments or political context rather than the bill's substantive content, thus introducing noise to the model.
Text standardization to ALMG format. Each document was reformatted according to the ALMG convention, concatenating the summary before the main body of the bill text to ensure structural consistency across documents.
After preprocessing, the dataset comprised 49,868 legislative bills from ALMG, including 3,754 similar pairs, and 71,898 bills from the Chamber of Deputies, with 43,484 similar pairs.
The positive pairs correspond to similarity links previously established by legislative analysts at ALMG and the Brazilian Chamber of Deputies during the regular legislative workflow. These links serve as the reference labels used throughout our experiments and reflect expert assessments that two bills address substantially the same legislative matter or pursue equivalent regulatory objectives. Consequently, the labels capture substantive legislative relatedness rather than purely lexical or semantic similarity.
4 ReSB2 Framework: Retrieving Similar Brazilian State Bills
This section describes ReSB2, the proposed framework for retrieving similar legislative bills 2. The framework is designed to assist legislative staff by automatically retrieving bills that are considered substantively related according to institutional legislative practice, thereby supporting analysts in establishing similarity links, improving both the efficiency and accuracy of the legislative review process. ReSB2 operates through four main stages, as illustrated in Figure 2: (1) model training; (2) embedding generation and indexing; (3) similarity search; and (4) explanation generation.
By combining domain-specific adaptation, task-focused fine-tuning and efficient vector-based retrieval aligned with institutional business rules, ReSB2 provides a robust and scalable solution for retrieving substantively related legislative bills through semantic representations in Brazilian state parliaments.
The similarity search process follows a two-stage architecture, combining a bi-encoder for large-scale candidate retrieval with a cross-encoder for fine-grained re-ranking. This architecture has become a standard and well-supported design pattern in modern Information Retrieval (IR) systems. Recent evaluations across multiple domains, including argument retrieval [36] and semantic textual relatedness [22], reinforce the effectiveness of this architecture by achieving an effective balance between scalability and ranking precision.
In the first stage, a bi-encoder [24] maps queries and documents into a shared embedding space, enabling efficient k-nearest neighbors (k-NN) search over large corpora. The second stage then refines the top-k candidates using a cross-encoder, which jointly encodes the query–document pair to capture token-level interactions. Cross-encoders consistently deliver substantial ranking improvements due to their richer attention mechanisms and deeper semantic alignment [5, 16].
Overall, the combination of bi-encoder retrieval followed by cross-encoder re-ranking is widely recognized as a state-of-the-art IR strategy, supported by contemporary empirical evidence and increasingly adopted in high-precision applications. Additionally, the framework supports interpretability through the application of Integrated Gradients (IG) [30], enabling users to visualize which textual elements most influenced the model's similarity predictions.
The base model used in ReSB2 is Multilingual ModernBERT (mmBERT) Base [20], which contains approximately 307 million parameters. This model was selected because it combines the efficiency of the ModernBERT architecture with multilingual support, including Portuguese.
4.1 Model training
The model training stage forms the foundation of the proposed framework, enabling language models to effectively capture the semantic and structural characteristics of legislative texts. This phase involves both continued pre-training and fine-tuning procedures, ensuring the models are adapted to the legal and legislative domain 3. Continued pre-training allows the base language model to internalise domain-specific vocabulary and stylistic patterns from legislative corpora, while fine-tuning optimises performance for the similarity-detection task. Together, these steps ensure that the resulting models can generate high-quality embeddings and accurate similarity judgments, forming the backbone of the retrieval process in subsequent stages.
4.1.1 Continued pre-training. Before fine-tuning, the base model underwent a continued pre-training phase using the Masked Language Modeling (MLM) objective. The goal was to expose the model to the syntactic patterns, terminology, and legal-drafting style typical of legislative texts, enabling it to better adapt to the linguistic characteristics of this domain. The resulting domain-adapted model was then used as the initialisation point for both the bi-encoder and cross-encoder fine-tuning stages.
4.1.2 Fine-tuning the Bi-Encoder. The next phase involved fine-tuning for embedding generation using a bi-encoder architecture, implemented following the Sentence-BERT (SBERT) framework [24]. In this setup, each legislative bill is independently encoded as a dense embedding vector, meaning the model processes one document per context window. The similarity between two bills is then computed by comparing their embeddings.
The main advantage of this approach is its efficiency for large-scale retrieval: the resulting embeddings can be indexed and queried in vector databases, enabling fast, scalable search across thousands of documents. However, since the bi-encoder encodes texts separately, it does not fully capture fine-grained token-level interactions between the two bills. To address this limitation, a second re-ranking stage is later applied using a cross-encoder model.
During training, we employed the Multiple Negatives Symmetric Ranking Loss [15]. This loss function requires only positive (i.e., similar bills) pairs, while treating all other examples within the same batch as implicit negatives. This setup allows the model to efficiently learn to bring embeddings of similar bills closer together while pushing apart those of dissimilar bills, without the need to manually construct negative pairs. This procedure is particularly effective when the objective is to build a vector index for similarity-based retrieval systems.
4.1.3 Fine-tuning the Cross-Encoder. Finally, a cross-encoder model was fine-tuned to re-rank the candidate pairs retrieved by the bi-encoder. Unlike the previous architecture, the cross-encoder processes each pair of bills jointly within the same context window. This allows the attention mechanism to model detailed token-to-token interactions between the texts, generally leading to more accurate similarity judgments.
The loss function used in this stage was Binary Cross-Entropy with Logits Loss [13], suitable for binary classification of similarity. Although the cross-encoder is computationally more expensive and not suitable for large-scale searches, it is highly effective for refining the ranking of the top candidate pairs retrieved in the first stage. In this sense, it acts as a high-precision re-ranking layer, valuable for identifying nuanced semantic relationships between texts that are not evident from embeddings alone.
4.2 Embedding Generation and Indexing
The bi-encoder model was employed to generate dense embeddings for all legislative bills in the dataset. Each embedding represents the semantic content of a bill in a high-dimensional vector space, enabling efficient similarity-based retrieval.
Once generated, these embeddings were indexed using an indexing system such as Elasticsearch4 or OpenSearch5. These systems provide efficient k-nearest neighbors (k-NN) search, enabling efficient retrieval of semantically similar bills. In addition, they operate as full-fledged databases capable of storing both embeddings and structured metadata, enabling domain-specific business rules to be applied directly during retrieval.
4.3 Similarity Search
The similarity search process combines two complementary stages: embedding-based retrieval with the bi-encoder model and re-ranking with the cross-encoder model.
In the first stage, the query bill is encoded into a dense vector by the bi-encoder and compared against pre-indexed embeddings using the indexing mechanism described in the previous section. This step efficiently retrieves the top-k most similar documents based on vector similarity. Acting as a fast pre-selection filter, it substantially reduces the search space while maintaining high recall.
During this stage, metadata attributes associated with each bill can also be used to impose constraints before the k-NN search, thereby integrating business rules directly into the retrieval pipeline. This hybrid approach, combining semantic retrieval with structured metadata filtering, ensures both context-aware search and alignment with institutional legislative requirements, thereby enhancing the practical applicability of the ReSB2 framework in real-world legislative environments.
In the second stage, each retrieved candidate bill is paired with the query text and jointly processed by the cross-encoder, which was fine-tuned for pairwise similarity classification. The cross-encoder generates a refined similarity score for each pair, allowing the retrieval results to be re-ranked by semantic relevance.
Overall, this two-stage retrieval strategy leverages the scalability of embedding-based retrieval and the precision of token-level interaction modeling, ensuring both efficiency and accuracy in retrieving candidate bills for substantive legislative similarity assessment.
4.4 Explanation generation
To provide additional insight into the cross-encoder stage, we apply Integrated Gradients (IG) [30], a widely adopted Explainable Artificial Intelligence (XAI) method for attributing model predictions to individual input features. IG estimates the contribution of each token in the input sequence processed by the cross-encoder (i.e., the two bills being evaluated for similarity), indicating how much each token positively or negatively influences the final similarity score. The contribution of the padding token is used as the baseline reference. Since IG generates token-level attributions, each word receives the average relevance of its subwords.
For visualization, we scale word contributions to [-1,1], keep only positive values above 0.2, and highlight them in shades of green: lighter tones show weaker influence, and darker ones indicate stronger relevance. This filtering and color scheme reduces visual clutter and helps users quickly identify the key words that guided the cross-encoder's decision.
Although transformer-based models remain largely opaque due to their complex internal mechanisms, IG provides a practical approach for obtaining local explanations of model predictions. Rather than revealing the complete reasoning process of the cross-encoder, IG highlights input features that contribute to a given prediction, offering partial transparency into similarity judgments. This information allows users to inspect which terms contributed most to the predicted relevance scores, particularly in cases where similar legislative bills share influential concepts.
5 Offline Evaluation
This section describes the experimental setup used to evaluate the proposed framework for identifying similar legislative bills. The experiments were designed to assess model performance under realistic legislative conditions, in line with the specific operational and institutional constraints of the Brazilian state legislature. All experiments used consistent training and evaluation procedures to ensure comparability, reproducibility, and statistical reliability.
5.1 Hyperparameters
The continued pre-training stage was executed for five epochs, while fine-tuning of both the bi-encoder and cross-encoder was performed for three epochs, balancing computational cost with sufficient training time to learn meaningful patterns.
All models were trained with a fixed hyperparameter configuration. Preliminary experiments indicated that this setup provided stable convergence and consistent validation loss. Training employed the AdamW optimizer with a learning rate of 5e-5, β1 = 0.9, β2 = 0.999, weight decay of 0.01, dropout probability of 0.1, and a warm-up ratio of 0.1. All other settings used the default parameters of the Hugging Face Transformers library. 6
To obtain robust performance estimates, each experiment was executed five times with different random seeds (148937, 529411, 785321, 903112, 662977), and the results were averaged.
5.2 Experimental Setup
The experiments were structured to reflect practical use cases of the legislative analysis process. Two search configurations were evaluated: (i) a session-restricted setting, in which similarity searches were limited to bills from the same legislative session, mirroring internal procedural rules that differentiate ongoing bills from those already concluded. (ii) a no-restriction setting, in which cross-session comparisons were allowed to capture broader patterns of similarity across the entire corpus.
The indexing system used was OpenSearch7, which provides efficient k-NN search. The retrieval pipeline followed the ReSB2 framework described in Section 4.3. We evaluated two variants: (i) the full two-stage pipeline, combining bi-encoder retrieval with cross-encoder re-ranking (with the first stage using k=50 in the k-NN search), and (ii) the first stage alone, consisting of the bi-encoder with indexing, to assess the impact of cross-encoder refinement on retrieval quality.
For evaluation, we use Recall to measure the proportion of relevant documents retrieved among the candidates. As most bills have only one relevant document, the task corresponds to a known-item search scenario [26]. In this setting, precision-based metrics are inherently limited, for example, Precision@10 has a maximum value of 0.1. Consequently, Mean Reciprocal Rank (MRR) is particularly suitable, as it captures how early the relevant document appears in the ranked list. Together, Recall and MRR provide complementary perspectives on retrieval effectiveness and the practical usefulness of ranked results in legislative similarity search.
5.3 Dataset Construction
To fine-tune the similarity models, we constructed pairs of legislative bills from SB2 Dataset. Pair sampling was designed to ensure diversity and class balance, using a 25%–75% ratio of similar to non-similar pairs. For ALMG bills, this resulted in 3,754 similar and 11,262 non-similar pairs; for the Chamber of Deputies bills, 43,484 similar and 130,444 non-similar pairs. Non-similar pairs were randomly sampled to avoid topical or temporal biases.
In the SB2 dataset, data from the Brazilian Chamber of Deputies were incorporated as an external corpus for knowledge transfer, enriching the linguistic and structural diversity available during training. The procedural rules described in Section 5.2 (e.g., restricting comparisons to bills within the same legislative session) can only be applied to the ALMG data, since legislative session metadata are not available for the Chamber of Deputies. For this reason, all evaluation was carried out exclusively on ALMG bills, ensuring methodological consistency and alignment with real-world operational conditions.
For training, all Chamber of Deputies pairs were combined with 50% of the ALMG pairs; from this combined set, 90% (163,343 pairs) were used for training and 10% (18,150 pairs) for validation. The remaining 50% of the ALMG pairs were reserved for testing, yielding a test set of 7,510 pairs, including 1,877 labeled as similar. From these 7,510 test pairs, 11,701 distinct bills were extracted and then indexed in the indexing system. The 1,877 similar pairs were used to construct the queries, formatted as (bill, list of similar bills), producing 1,795 queries under the same-session restriction and 3,155 under the unrestricted configuration.
Bills not selected for pair construction (either for training or testing) were repurposed for continued pre-training using the Masked Language Modelling (MLM). The amount of data available for MLM varied across the five experimental repetitions due to the randomness of non-similar pair sampling, but averaged approximately 42,000 documents.
Figure 3 shows the distribution of token counts in ALMG bill texts after preprocessing, with several documents exceeding mmBERT's 8,192-token context limit. A 1,500-token window (≈ 96th percentile) was chosen to fully capture most bills. Accordingly, bi-encoder training and continued pre-training used a maximum sequence length of 1,500 tokens, while the cross-encoder, which processes both bills jointly, employed a 3,000-token window (≈ 1,500 tokens per bill) to ensure consistency across stages.
5.4 Baselines
We compare ReSB2 against a diverse set of baseline models, covering both domain-specific and general-purpose approaches:
(i) JurisBERT [32], a transformer model trained from scratch on Brazilian Portuguese legal texts and subsequently fine-tuned for the Semantic Textual Similarity (STS) task.
(ii) BERTikal [23], a domain-adapted transformer model based on BERTimbau [29], further pre-trained on Brazilian Portuguese legal corpora to better capture domain-specific language patterns.
Both JurisBERT and BERTikal are tailored to legal language in Brazilian Portuguese. However, as they are built upon the original BERT architecture, they are constrained to a maximum context window of 512 tokens.
(iii) OpenAI Text Embeddings 3 Large [21], a general-purpose multilingual embedding model that captures broad semantic relationships across domains and languages. This model supports a significantly larger context window (up to 8192 tokens), providing a strong cross-domain baseline.
(iv) mmBERT [20], a multilingual ModernBERT model, serves as the base architecture for ReSB2. This baseline allows us to isolate and measure the impact of the domain-adaptive pre-training (MLM) and fine-tuning strategies introduced in our framework. For a fair comparison, mmBERT is evaluated using the same context window as ReSB2 (1500 tokens in the bi-encoder).
(v) ReSB2 ALMG, a variant of our framework trained exclusively on data from the Legislative Assembly of Minas Gerais (ALMG). As described in Section 5.3, the full SB2 training set includes legislative data from both ALMG and the Brazilian Chamber of Deputies, the latter serving as an external corpus for knowledge transfer, while all evaluations are conducted on ALMG bills. This baseline assesses the contribution of incorporating external legislative data during training. The same context windows as ReSB2 are used (1500 tokens for the bi-encoder and 3000 tokens for the cross-encoder).
Together, these baselines offer complementary perspectives for evaluation. OpenAI embeddings provide insight into general-purpose, cross-domain performance with extended context capacity. JurisBERT and BERTikal establish strong legal-domain baselines, although they are constrained by the limited context window of the original BERT architecture. Finally, mmBERT (the base model used in ReSB2) and ReSB2 ALMG (a variant trained on a reduced dataset) enable a controlled analysis of the impact of domain-specific training under different data compositions, while maintaining the same architecture and context window as the full ReSB2 framework.
5.5 Results
The results, summarized in Table 1 and visualized in Figure 4, reveal substantial performance differences across models and evaluation settings, highlighting the challenges of legislative similarity retrieval and the strengths of the proposed ReSB2.
Table 1: Retrieval performance across different models and experimental conditions
Recall@n | MRR@n | |||||||
|---|---|---|---|---|---|---|---|---|
n=5 | n=10 | n=15 | n=20 | n=5 | n=10 | n=15 | n=20 | |
Same legislative session | ||||||||
JurisBERT | 0.571 ± 0.010 | 0.676 ± 0.006 | 0.729 ± 0.006 | 0.766 ± 0.003 | 0.421 ± 0.005 | 0.435 ± 0.005 | 0.440 ± 0.005 | 0.442 ± 0.005 |
BERTikal | 0.481 ± 0.003 | 0.553 ± 0.005 | 0.598 ± 0.005 | 0.635 ± 0.006 | 0.366 ± 0.006 | 0.376 ± 0.005 | 0.379 ± 0.005 | 0.381 ± 0.005 |
OpenAI Emb Large | 0.803 ± 0.014 | 0.864 ± 0.008 | 0.885 ± 0.009 | 0.903 ± 0.008 | 0.655 ± 0.009 | 0.663 ± 0.008 | 0.665 ± 0.008 | 0.666 ± 0.008 |
mmBERT | 0.416 ± 0.012 | 0.477 ± 0.014 | 0.522 ± 0.015 | 0.555 ± 0.009 | 0.320 ± 0.011 | 0.329 ± 0.010 | 0.332 ± 0.010 | 0.334 ± 0.010 |
ReSB2 ALMG | 0.777 ± 0.013 | 0.868 ± 0.011 | 0.912 ± 0.013 | 0.939 ± 0.010 | 0.601 ± 0.005 | 0.613 ± 0.005 | 0.616 ± 0.005 | 0.618 ± 0.005 |
ReSB2 stage 1 | 0.851 ± 0.008 | 0.905 ± 0.007 | 0.929 ± 0.005 | 0.942 ± 0.005 | 0.683 ± 0.005 | 0.690 ± 0.005 | 0.692 ± 0.005 | 0.693 ± 0.005 |
ReSB2 | 0.856 ± 0.008 | 0.913 ± 0.008 | 0.937 ± 0.006 | 0.951 ± 0.005 | 0.686 ± 0.010 | 0.694 ± 0.010 | 0.696 ± 0.010 | 0.697 ± 0.009 |
No restriction | ||||||||
JurisBERT | 0.427 ± 0.004 | 0.486 ± 0.003 | 0.524 ± 0.004 | 0.550 ± 0.006 | 0.331 ± 0.007 | 0.339 ± 0.006 | 0.342 ± 0.007 | 0.344 ± 0.006 |
BERTikal | 0.419 ± 0.007 | 0.465 ± 0.008 | 0.490 ± 0.007 | 0.510 ± 0.007 | 0.342 ± 0.004 | 0.348 ± 0.004 | 0.350 ± 0.004 | 0.351 ± 0.004 |
OpenAI Emb Large | 0.724 ± 0.004 | 0.793 ± 0.005 | 0.821 ± 0.004 | 0.838 ± 0.005 | 0.568 ± 0.007 | 0.578 ± 0.007 | 0.580 ± 0.007 | 0.581 ± 0.007 |
mmBERT | 0.391 ± 0.004 | 0.427 ± 0.004 | 0.448 ± 0.003 | 0.465 ± 0.004 | 0.324 ± 0.002 | 0.329 ± 0.002 | 0.330 ± 0.002 | 0.331 ± 0.002 |
ReSB2 ALMG | 0.685 ± 0.017 | 0.793 ± 0.012 | 0.850 ± 0.005 | 0.882 ± 0.002 | 0.508 ± 0.023 | 0.523 ± 0.022 | 0.527 ± 0.022 | 0.529 ± 0.021 |
ReSB2 stage 1 | 0.735 ± 0.011 | 0.811 ± 0.006 | 0.846 ± 0.004 | 0.868 ± 0.002 | 0.571 ± 0.009 | 0.582 ± 0.009 | 0.584 ± 0.009 | 0.586 ± 0.009 |
ReSB2 | 0.724 ± 0.029 | 0.821 ± 0.012 | 0.868 ± 0.004 | 0.892 ± 0.002 | 0.529 ± 0.045 | 0.542 ± 0.042 | 0.546 ± 0.042 | 0.547 ± 0.042 |
mmBERT achieves the lowest performance across all evaluated metrics, which is expected given that it does not incorporate any domain-specific adaptation to the legal context. BERTikal performs better, as it is pre-trained on Brazilian legal corpora and can therefore capture domain-relevant linguistic patterns. JurisBERT further improves upon these results, as it is not only trained on legal texts but also fine-tuned for the Semantic Textual Similarity (STS) task, making it particularly well-suited for this evaluation setting.
OpenAI Text Embeddings 3 Large also operates without task-specific fine-tuning in our experiments. However, its substantially deeper architecture, larger context window, and broader training corpus result in a more expressive embedding space. Despite being a general-purpose model, its competitive performance suggests that it captures legislative or formal-text patterns more effectively, enabling stronger semantic matching.
Given the nature of the task, retrieving semantically similar bills to support legislative analysis, it can be characterized as a high recall system, making Recall a more critical metric than MRR. In practice, analysts are willing to examine all documents retrieved within a reasonable cutoff to ensure that no redundant bills advance through the legislative process. Under these conditions, maximizing Recall becomes the primary objective.
Since the Recall values of ReSB2, ReSB2 Stage 1, ReSB2 ALMG, and OpenAI Text Embeddings 3 Large fall within overlapping margins of error (mean ± standard deviation), we further assess statistical significance using the Wilcoxon signed-rank test, which is appropriate for paired samples. Specifically, we compare each of ReSB2 Stage 1, ReSB2 ALMG, and OpenAI Text Embeddings 3 Large against ReSB2, testing whether the median difference in Recall is zero at a significance level of α = 0.05. This analysis is conducted for both the same legislative session and the no-restriction settings.
ReSB2 ALMG achieves strong performance in terms of Recall@10 and Recall@20, remaining competitive with the full ReSB2 model. However, the Recall values of ReSB2 are statistically significantly higher than those of ReSB2 ALMG in both the same legislative session and no-restriction settings (Wilcoxon signed-rank test, W = 15.0, p ≤ 0.05 across all Recall cutoffs). ReSB2 ALMG also exhibits considerably lower MRR values, indicating that although relevant documents are often retrieved within the top-k results, they tend to appear in lower-ranked positions.
The substantial performance gap between ReSB2 and mmBERT highlights the effectiveness of the training strategy in Section 4. Furthermore, the significant difference between ReSB2 and ReSB2 ALMG evidences the benefit of incorporating the Chamber of Deputies’ external knowledge base during training.
When restricting the search space to bills within the same legislative session, all models naturally achieve higher Recall and MRR, reflecting the more homogeneous and temporally aligned nature of these documents. Across all metrics and values of n, the full ReSB2 pipeline consistently outperforms all baselines (Wilcoxon signed-rank test, W = 15.0, p ≤ 0.05 across all Recall cutoffs). These improvements are pronounced at higher cutoffs, for instance, ReSB2 reaches 0.951 Recall@20 in the same-session scenario, surpassing all baselines by a considerable margin. Gains from ReSB2 Stage 1 are also evident, indicating that even without the re-ranking, the specialized retrieval stage substantially enhances performance.
In the no-restriction scenario, where candidate bills come from all sessions and no metadata filters are applied, performance declines across all models due to increased heterogeneity. Even under these more challenging conditions, ReSB2 maintains superior performance compared to ReSB2 ALMG, as previously discussed. It also outperforms OpenAI Text Embeddings 3 Large, with statistically significant gains (Wilcoxon signed-rank test, W = 8.0, p ≤ 0.05 for Recall@5, and W = 15.0, p ≤ 0.05 for the remaining Recall cutoffs).
While Stage 1 achieves slightly better results at the very top of the ranking compared to ReSB2 (e.g., Recall@5, with W = 15.0, p ≤ 0.05), the full ReSB2 framework demonstrates more robust performance as n increases. For Recall@10, although ReSB2 shows numerically higher values than Stage 1, the difference is not statistically significant (W = 12.0, p > 0.05). However, for Recall@15 and Recall@20, ReSB2 significantly outperforms Stage 1 (W = 15.0, p ≤ 0.05). This pattern suggests that, in more heterogeneous settings, the cross-encoder introduces a re-ranking effect that may slightly alter the top-ranked positions, but becomes increasingly beneficial at deeper ranks by correcting mid-ranking errors and improving the retrieval of relevant bills beyond the first positions. This behavior also helps explain the higher variance observed in the full ReSB2 framework under these conditions.
As discussed earlier, Recall is the most critical metric for this task, and the full ReSB2 framework consistently achieves the best performance under this criterion, making it the most reliable approach for real-world legislative retrieval scenarios.
6 Human-Centric Evaluation
This section presents a human-centric evaluation of ReSB2 in a real-world production environment. Rather than relying solely on offline experiments, we assess the framework as a tool for human augmentation within the legislative workflow of the Minas Gerais Parliament (ALMG). The system implementing ReSB2 has been deployed for four months; however, consistent usage began only two months ago. Since then, a total of 2,319 queries have been issued by 11 distinct users.
6.1 Professional Feedback
When a bill is queried, the system presents the top-10 retrieved bills as candidates for linkage. Users can provide explicit feedback for each suggested bill through a binary signal (positive or negative), indicating whether the retrieved bill is similar to the queried one. If no relevant bill is retrieved, users may manually specify the appropriate linkage. Importantly, this feedback does not constitute the final linking decision; rather, it serves as optional input to support future model improvements.
In total, 136 explicit feedback instances were collected, covering 90 distinct bills. Among these, 25 bills were excluded from the analysis, as they received neither positive feedback nor manual linkage, making it impossible to determine whether no similar bill existed or whether it was not retrieved by ReSB2. This results in a final set of 65 evaluated bills.
Out of these, 53 bills had at least one correct suggestion among the results presented by ReSB2, as indicated by positive user feedback, while in 12 cases the correct bill had to be manually inserted. This corresponds to a success rate of 81.5%. Given that users are presented with the top-10 results, the corresponding metrics are Recall@10 = 0.815 and MRR@10 = 0.619.
6.2 Field Study
A complementary evaluation was conducted based on actual linkage decisions made by users. In total, 197 bills were linked following a query and the corresponding retrieval of candidate bills by ReSB2. Among these, 162 cases included the final linked bill within the top-10 suggestions provided by the system, resulting in a success rate of 82.2%. Under this setting, the observed performance corresponds to Recall@10 = 0.822 and MRR@10 = 0.545.
6.3 Qualitative Analysis
Beyond quantitative results, human-centric evaluation revealed a fundamental trade-off in legislative retrieval: balancing large context windows with domain-specific expertise.
The Challenge of Scale (Fiscal Links). In cases involving complex fiscal legislation, particularly bills containing extensive annexes, tables, and structured forms, documents tend to be significantly longer. We observed that the general-purpose model OpenAI Text Embeddings 3 Large, which support context windows of up to 8,152 tokens, occasionally outperformed the base ReSB2 architecture. These cases typically occurred when the relevant semantic signal was located near the end of long documents or embedded within technical attachments. In such scenarios, the ability to process larger contexts allowed generalist models to capture information that would otherwise be truncated or diluted in models constrained by shorter input limits.
The Advantage of Domain Knowledge (Procedural Links). Conversely, ReSB2 consistently demonstrated superior performance in identifying procedural and institution-specific relationships that general-purpose models failed to capture. These connections often arise from implicit legislative rules and workflow conventions rather than surface-level textual similarity. For instance, the framework successfully linked bills with distinct vocabularies but equivalent procedural roles or legislative statuses. This behavior highlights the importance of domain-adaptive training, enabling the model to bridge structurally related yet lexically divergent documents, an ability that general-purpose models typically lack.
6.4 Discussion
In both the Professional Feedback and Field Study analyses, it is not possible to determine whether procedural filtering rules were applied; consequently, it is unclear whether the queries correspond to the same legislative session or to a no-restriction scenario. Nevertheless, the observed Recall@10 and MRR@10 values are consistent with those reported in Table 4 under the no-restriction setting. This alignment suggests that the performance observed in offline experiments generalizes well to real-world usage.
The interaction of domain experts with the system demonstrates that ReSB2 effectively supports the identification of relevant legislative connections. Moreover, the feedback collected during the qualitative analysis provides valuable insights that can guide future improvements to the framework.
In practice, ReSB2 shifts the interaction between analysts and legislative documents from manually searching large collections to reviewing a ranked list of candidate bills. This human-in-the-loop workflow preserves expert judgment while focusing the analyst's attention on the most relevant candidates. Although our evaluation did not directly measure productivity gains or task completion times, the sustained use of the system (2,319 queries issued by 11 users) indicates that it integrates naturally into the legislative analysis process.
7 Explanation example
Figure 5 presents an example of a local explanation generated by ReSB2 using Integrated Gradients (IG). Green highlights indicate positive contributions to the predicted similarity, with darker shades representing stronger contributions of individual words.
In the example, the highlighted words correspond to the core regulatory concepts shared by the two bills, such as the prohibition of associating gifts with the sale of food products. Despite differences in wording, the explanation shows that the model focuses on legally meaningful concepts rather than relying solely on lexical overlap. These explanations provide legislative analysts with additional evidence to verify whether a suggested linkage is supported by relevant legislative content, thereby increasing transparency and facilitating human validation within the retrieval process.
8 Conclusions and Future Work
This work presented ReSB2, a framework for retrieving and linking similar legislative bills in the Brazilian legislative context. The framework supports human–machine collaboration and helps reduce redundancy in the lawmaking process.
Compared to OpenAI's embedding-based solution, ReSB2 achieved superior performance across all evaluated metrics. Moreover, the use of a local model enables the application of interpretability methods, such as Integrated Gradients, allowing analysts to inspect word-level contributions to similarity judgments and providing partial insight into the re-ranking process. These results demonstrate that ReSB2 not only achieves strong performance but also provides a cost-effective, domain-adapted, and partially transparent solution suitable for integration into legislative workflows.
Domain specialists interacted with the link suggestions generated by ReSB2, achieving a success rate above 80%. This human-in-the-loop evaluation indicates that the framework effectively complements legal expertise, acting as an assistive tool that surfaces relevant candidates within a large and complex search space. Furthermore, because the architecture combines generic neural retrieval components with institution-specific training data and procedural filters, it can potentially be adapted to legislatures in other countries, provided equivalent corpora and similarity annotations are available.
Limitations. A key limitation of the dataset is the presence of semantically similar bills that were never formally linked due to political factors, resulting in structural false negatives that may affect evaluation. In addition, some legislative bills exceed the maximum input length supported by the current model configuration, requiring truncation that may omit relevant long-range information.
Future Work. Future work will focus on incorporating a user feedback loop to iteratively refine similarity signals based on expert input, thereby mitigating false negatives. Another promising direction is to investigate architectures and training strategies capable of processing longer legislative documents.
Acknowledgments. This work was supported by Legislative Assembly of Minas Gerais (Assembleia Legislativa de Minas Gerais, ALMG) and by CNPq, CAPES, FAPEMIG.
Declaration of Generative AI use. Generative AI was used only for linguistic refinement and grammatical editing.
Notes
1Data available at https://huggingface.co/datasets/lucas-lage/SB2-Dataset.
2Framework code available at https://github.com/LucasLage/ReSB2-Framework.
3Trained models are publicly available:
Source
References
[1] Giuseppe Abrami, Daniel Bundan, Chrisowaladis Manolis, and Alexander Mehler. 2025. VR-ParlExplorer: A Hypertext System for the Collaborative Interaction in Parliamentary Debate Spaces. In Proceedings of the 36th ACM Conference on Hypertext and Social Media(HT ’25). Association for Computing Machinery, New York, NY, USA, 177–183. https://doi.org/10.1145/3720553.3746672 — ACM HyperText copy
[2] Maristella Agosti, Roberto Colotti, and Girolamo Gradenigo. 1991. A two-level hypertext retrieval model for legal data. In Proceedings of the 14th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Chicago, Illinois, USA) (SIGIR ’91). Association for Computing Machinery, New York, NY, USA, 316–325. https://doi.org/10.1145/122860.122892
[3] Hidelberg O. Albuquerque, Rosimeire Costa, Gabriel Silvestre, Ellen Souza, Nádia F. F. da Silva, Douglas Vitório, Gyovana Moriyama, Lucas Martins, Luiza Soezima, Augusto Nunes, Felipe Siqueira, João P. Tarrega, Joao V. Beinotti, Marcio Dias, Matheus Silva, Miguel Gardini, Vinicius Silva, André C. P. L. F. de Carvalho, and Adriano L. I. Oliveira. 2022. UlyssesNER-Br: A Corpus of Brazilian Legislative Documents for Named Entity Recognition. In Computational Processing of the Portuguese Language: 15th International Conference, PROPOR 2022, Fortaleza, Brazil, March 21–23, 2022, Proceedings (Fortaleza, Brazil). Springer-Verlag, Berlin, Heidelberg, 3–14. https://doi.org/10.1007/978-3-030-98305-51
[4] Gisliany Alves, Breno Santana Santos, Marianne Silva, and Ivanovitch Silva. 2025. Brazilian Portuguese Legislative Documents: A Dataset from the Legislative Assembly of Rio Grande do Norte. Mendeley Data. https://doi.org/10.17632/9df7zthpfd.1
[5] Amir Askari, Abbas Abolghasemi, Gabriella Pasi, et al. 2024. Injecting the Score of the First-Stage Retriever as Text Improves BERT-Based Re-Rankers. Discover Computing 27, 15 (2024). https://doi.org/10.1007/s10791-024-09435-8
[6] Assembleia Legislativa do Estado de Minas Gerais (Brasil - Minas Gerais). 2025. Regimento Interno da Assembleia Legislativa do Estado de Minas Gerais – texto atualizado até a Deliberação da Mesa nº2.860, de 14 abr. 2025. https://www.almg.gov.br/atividade-parlamentar/leis/legislacao-mineira/lei/texto/?tipo=RAL&num=5176&ano=1997&comp=&cons=1. Accessed: 2026-04-13.
[7] Daniele Bertillo, Andrea de Donato, Carlo Marchetti, and Paolo Merialdo. 2023. Enhancing Accessibility of Parliamentary Video Streams: AI-Based Automatic Indexing Using Verbatim Reports. In Proceedings of the 1st Legal Information Retrieval meets Artificial Intelligence Workshop (LIRAI 2023) co-located with the 34th ACM Hypertext Conference (HT 2023)(CEUR Workshop Proceedings, Vol. 3594), Sabine Wehnert, Manuel Fiorelli, Davide Picca, Ernesto William De Luca, and Armando Stellato (Eds.). CEUR-WS.org, Rome, Italy, 31–40. https://ceur-ws.org/Vol-3594/paper2.pdf
[8] Wojciech Cyrul and Tomasz Pełech-Pilichowski. 2020. Legislating in hypertext. The Opole Studies in Administration and Law 18, 2 (Oct. 2020), 27–42. https://doi.org/10.25167/osap.2178
[9] Câmara dos Deputados (Brasil). 2025. Regimento Interno da Câmara dos Deputados – texto atualizado até a Resolução nº16, de 2025. https://www2.camara.leg.br/atividade-legislativa/legislacao/regimento-interno-da-camara-dos-deputados. Accessed: 2026-04-13.
[10] Ernesto William De Luca, Manuel Fiorelli, Davide Picca, Armando Stellato, and Sabine Wehnert. 2023. Legal Information Retrieval meets Artificial Intelligence (LIRAI). In Proceedings of the 34th ACM Conference on Hypertext and Social Media (Rome, Italy) (HT ’23). Association for Computing Machinery, New York, NY, USA, Article 46, 4 pages. https://doi.org/10.1145/3603163.3610575 — ACM HyperText copy
[11] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT. Association for Computational Linguistics, 4171–4186. https://doi.org/10.18653/v1/n19-1423
[12] Shangsheng Gao, Li Gao, Qi Li, and Jianjun Xu. 2023. Application of large language model in intelligent Q&A of digital government. In Proceedings of the 2023 2nd International Conference on Networks, Communications and Information Technology (Qinghai, China) (CNCIT ’23). Association for Computing Machinery, New York, NY, USA, 24–27. https://doi.org/10.1145/3605801.3605806
[13] Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. MIT Press. http://www.deeplearningbook.org.
[14] Graham Greenleaf, Andrew Mowbray, and Philip Chung. 2018. Building sustainable free legal advisory systems: Experiences from the history of AI & law. Computer Law & Security Review 34, 2 (2018), 314–326. https://doi.org/10.1016/j.clsr.2018.02.007
[15] Matthew Henderson, Rami Al-Rfou, Brian Strope, Yun hsuan Sung, László Lukács, Ruiqi Guo, Sanjiv Kumar, Bálint Miklós, and Ray Kurzweil. 2017. Efficient Natural Language Response Suggestion for Smart Reply. arxiv:1705.00652 [cs.CL]
[16] Meng Lu, Catherine Chen, and Carsten Eickhoff. 2025. Pathway to Relevance: How Cross-Encoders Implement a Semantic Variant of BM25. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (Eds.). Association for Computational Linguistics, Suzhou, China, 25536–25558. https://doi.org/10.18653/v1/2025.emnlp-main.1297
[17] Masahiro Makino, Yuya Asazuma, Shota Sasaki, and Jun Suzuki. 2024. The Impact of Integration Step on Integrated Gradients. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, Neele Falk, Sara Papi, and Mike Zhang (Eds.). Association for Computational Linguistics, St. Julian's, Malta, 279–289. https://doi.org/10.18653/v1/2024.eacl-srw.22
[18] Marios Evangelos Mamalis, Evangelos Kalampokis, Areti Karamanou, Petros Brimos, and Konstantinos Tarabanis. 2024. Can Large Language Models Revolutionalize Open Government Data Portals? A Case of Using ChatGPT in statistics.gov.scot. In Proceedings of the 27th Pan-Hellenic Conference on Progress in Computing and Informatics (Lamia, Greece) (PCI ’23). Association for Computing Machinery, New York, NY, USA, 53–59. https://doi.org/10.1145/3635059.3635068
[19] Juan Marciano, Vinicius Machado, and Arlino Araújo. 2025. LegisPL-BR - Dataset de projetos de leis brasileiros. In Anais do VII Dataset Showcase Workshop (Fortaleza/CE). SBC, Porto Alegre, RS, Brasil, 58–70. https://doi.org/10.5753/dsw.2025.247728
[20] Marc Marone, Orion Weller, William Fleshman, Eugene Yang, Dawn Lawrie, and Benjamin Van Durme. 2025. mmBERT: A Modern Multilingual Encoder with Annealed Language Learning. arxiv:2509.06888 [cs.CL] https://arxiv.org/abs/2509.06888
[21] OpenAI. 2025. text-embedding-3-large. https://platform.openai.com/docs/models/text-embedding-3-large. Accessed: 2026-04-13.
[22] Jesus German Ortiz Barajas, Gemma Bel-enguix, and Helena Goméz-adorno. 2024. MBZUAI-UNAM at SemEval-2024 Task 1: Sentence-CROBI, a Simple Cross-Bi-Encoder-Based Neural Network Architecture for Semantic Textual Relatedness. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024), Atul Kr. Ojha, A. Seza Doğruöz, Harish Tayyar Madabushi, Giovanni Da San Martino, Sara Rosenthal, and Aiala Rosá (Eds.). Association for Computational Linguistics, Mexico City, Mexico, 1071–1079. https://doi.org/10.18653/v1/2024.semeval-1.155
[23] Felipe Maia Polo, Gabriel Caiaffa Floriano Mendonça, Kauê Capellato J Parreira, Lucka Gianvechio, Peterson Cordeiro, Jonathan Batista Ferreira, Leticia Maria Paz de Lima, Antônio Carlos do Amaral Maia, and Renato Vicente. 2021. LegalNLP-Natural Language Processing methods for the Brazilian Legal Language. In Anais do XVIII Encontro Nacional de Inteligência Artificial e Computacional. SBC, 763–774.
[24] Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics. https://arxiv.org/abs/1908.10084
[25] Jacques Savoy. 1993. Searching information in legal hypertext systems. Artificial Intelligence and Law 2, 3 (1993), 205–232. https://doi.org/10.1007/BF00871890
[26] Hinrich Schütze, Christopher D Manning, and Prabhakar Raghavan. 2008. Introduction to information retrieval. Vol. 39. Cambridge University Press Cambridge.
[27] Mariana Silva, Gabriel Oliveira, Lucas Costa, and Gisele Pappa. 2024. Evaluating Domain-adapted Language Models for Governmental Text Classification Tasks in Portuguese. In Anais do XXXIX Simpósio Brasileiro de Bancos de Dados (Florianópolis/SC). SBC, Porto Alegre, RS, Brasil, 247–259. https://doi.org/10.5753/sbbd.2024.240508
[28] Mariana O. Silva, Gabriel P. Oliveira, Lucas G. L. Costa, and Gisele L. Pappa. 2025. GovBERT-BR: A BERT-Based Language Model for Brazilian Portuguese Governmental Data. In Intelligent Systems, Aline Paes and Filipe A. N. Verri (Eds.). Springer Nature Switzerland, Cham, 19–32. https://link.springer.com/chapter/10.1007/978-3-031-79032-42
[29] Fábio Souza, Rodrigo Nogueira, and Roberto Lotufo. 2020. BERTimbau: pretrained BERT models for Brazilian Portuguese. In 9th Brazilian Conference on Intelligent Systems, BRACIS, Rio Grande do Sul, Brazil, October 20-23 (to appear). https://dl.acm.org/doi/10.1007/978-3-030-61377-828
[30] Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017. Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning - Volume 70 (Sydney, NSW, Australia) (ICML’17). JMLR.org, 3319–3328. https://dl.acm.org/doi/10.5555/3305890.3306024
[31] S. Talebi, E. Tong, A. Li, et al. 2024. Exploring the performance and explainability of fine-tuned BERT models for neuroradiology protocol assignment. BMC Medical Informatics and Decision Making 24, 40 (2024). https://doi.org/10.1186/s12911-024-02444-z
[32] Charles F. O. Viegas, Bruno C. Costa, and Renato P. Ishii. 2023. JurisBERT: A New Approach that Converts a Classification Corpus into an STS One. In Computational Science and Its Applications – ICCSA 2023. Springer Nature Switzerland, 349–365. https://doi.org/10.1007/978-3-031-36805-924
[33] Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Griffin Thomas Adams, Jeremy Howard, and Iacopo Poli. 2025. Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 2526–2547. https://doi.org/10.18653/v1/2025.acl-long.127
[34] Sabine Wehnert, Manuel Fiorelli, Davide Picca, Ernesto William De Luca, and Armando Stellato. 2024. LIRAI’24: 2nd Workshop on Legal Information Retrieval meets Artificial Intelligence. In Proceedings of the 35th ACM Conference on Hypertext and Social Media (Poznan, Poland) (HT ’24). Association for Computing Machinery, New York, NY, USA, 390–392. https://doi.org/10.1145/3648188.3675120 — ACM HyperText copy
[35] Eve Wilson. 1992. Links and structures in hypertext databases for law. Cambridge University Press, USA, 194–211.
[36] Leixin Zhang and Daniel Braun. 2024. Twente-BMS-NLP at PerspectiveArg 2024: Combining Bi-Encoder and Cross-Encoder for Argument Retrieval. In Proceedings of the 11th Workshop on Argument Mining (ArgMining 2024), Yamen Ajjour, Roy Bar-Haim, Roxanne El Baff, Zhexiong Liu, and Gabriella Skitalinskaya (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 164–168. https://doi.org/10.18653/v1/2024.argmining-1.17
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime