<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Work on Sina's Page</title><link>https://sina.page/work/</link><description>Recent content in Work on Sina's Page</description><generator>Hugo</generator><language>en</language><copyright/><atom:link href="https://sina.page/work/index.xml" rel="self" type="application/rss+xml"/><item><title>Cell2Sentence</title><link>https://sina.page/work/cell2sentence/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://sina.page/work/cell2sentence/</guid><description>&lt;p&gt;Cell2Sentence explores how biological measurements can be represented as sequences that language models can process.&lt;/p&gt;
&lt;p&gt;The project adapts LLM ideas to single-cell biology, connecting representation learning, natural language modeling, and biological discovery.&lt;/p&gt;
&lt;p&gt;See the paper and code for technical details.&lt;/p&gt;</description></item><item><title>Machine Learning for Protein Biology</title><link>https://sina.page/work/protein-ml/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://sina.page/work/protein-ml/</guid><description>&lt;p&gt;During my PhD, I worked on machine-learning methods that infer biological and structural properties directly from protein sequence.&lt;/p&gt;
&lt;p&gt;The work covered several related problems: predicting intrinsically disordered regions and their functions, DNA-binding residues, protein-binding residues, crystal-structure quality, and sequence-derived signals associated with druggability.&lt;/p&gt;
&lt;p&gt;A recurring theme was &lt;strong&gt;learning useful representations from sequence when direct structural evidence is sparse or unavailable&lt;/strong&gt;. That meant thinking carefully about labels, evaluation, biological priors, and the gap between benchmark performance and useful scientific inference.&lt;/p&gt;</description></item><item><title>Patient Record Linkage at Scale</title><link>https://sina.page/work/patient-record-linkage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://sina.page/work/patient-record-linkage/</guid><description>&lt;h2 id="overview" class="relative group"&gt;Overview &lt;span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"&gt;&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#overview" aria-label="Anchor"&gt;#&lt;/a&gt;&lt;/span&gt;&lt;/h2&gt;&lt;p&gt;I designed and built a large-scale patient matching system combining learned representations with deterministic constraints.&lt;/p&gt;
&lt;p&gt;The system used neural embeddings, blocking strategies, clustering, and precision/recall evaluation techniques to link records across healthcare datasets while balancing accuracy and scalability.&lt;/p&gt;
&lt;h2 id="why-it-matters" class="relative group"&gt;Why it matters &lt;span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"&gt;&lt;a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#why-it-matters" aria-label="Anchor"&gt;#&lt;/a&gt;&lt;/span&gt;&lt;/h2&gt;&lt;p&gt;Entity resolution is a foundation problem for healthcare data. Small errors have large downstream effects, and the system must combine statistical learning with domain knowledge.&lt;/p&gt;</description></item></channel></rss>